The extended test plan separates three layers: application-visible integration, provider-facing callbacks, and documentary comparison. This separation makes it easier to identify what each result means. Integration tests demonstrate that a normal OpenSSL consumer can use the module. Callback tests reach interface boundaries that may not be directly controllable from a high-level consumer. Documentation establishes the intended meanings of the interfaces and version-specific configuration options.
The new integration harness generates each byte using a simple arithmetic pattern, reducing the risk that all tested messages contain only a single repeated value. For an input of length n, byte i is computed from (17i + 31) modulo 256. This is not a source of cryptographic randomness and is not presented as one. It is a deterministic way to exercise a variety of byte values while making every test message reproducible without storing separate binary fixtures.
The selected lengths cover ordinary small messages, output-size landmarks, SHA-256 block and padding boundaries, and larger application sizes. The set is 0, 1, 2, 3, 31, 32, 33, 55, 56, 57, 63, 64, 65, 119, 120, 121, 127, 128, 129, 255, 256, 257, 4095, 4096, 4097, 65535, 65536, and 65537. These values are not a proof of exhaustive coverage. They are a deliberate partition of boundary conditions that is more informative than choosing several arbitrary lengths.
SHA-256 operates on 64-byte message blocks, but padding and encoded message length also consume space in the final block. Values around 55 and 56 bytes, and corresponding positions in a later block, are therefore useful reference-test boundaries. The provider itself does not implement that padding; its backend does. Testing these values mainly establishes that the bridge passes lengths and bytes correctly and does not accidentally truncate at a boundary. The underlying algorithmic interpretation comes from the hash standard. [13]
The sizes around 4096 and 65536 bytes serve a different purpose. They resemble common application buffer landmarks, without claiming any special behavior of this host’s I/O implementation. A successful command-line comparison at those lengths shows that the tested path handles the full input. It does not demonstrate a particular buffering strategy inside the executable or establish that a single update callback received the entire message.
For each generated input, Python computes an expected digest through hashlib.sha256. The command-line operation is asked for binary output, and the harness compares bytes rather than a human-readable label. It checks the process exit code as well as the digest. This avoids accepting an empty output or a formatted error message as though it were a digest result.
The reference is a practical integration oracle, but it is not necessarily algorithmically independent. Python may use OpenSSL internally, and the provider bridge certainly does. The study therefore distinguishes a reference computation from an independent cryptographic implementation. Published known-answer values and a separate backend would strengthen an algorithm-validation argument, but the current work is principally about interface behavior.
The callback harness loads edu.so with the platform dynamic loader, obtains the initialization symbol, and retrieves the returned dispatch tables. It uses the public typed extraction helpers to obtain callback pointers. Because this specific module ignores the incoming core handle and callbacks, the harness can supply nulls for those arguments. A provider that depends on core allocation, error, or configuration callbacks would require a more complete test environment.
This is an important limitation of the method, not a hidden shortcut. The harness is white-box testing for a known implementation. Its purpose is to evaluate the digest guard, metadata, unsupported-operation response, and reinitialization behavior of that implementation. The ordinary integration tests still establish that the same module works when OpenSSL itself supplies the surrounding lifecycle.
The extension contains 37 integration cases: 28 message-length cases, three property cases, four activation cases, one missing-module case, and one parallel-process case. The callback program reports 3,362 assertions. These numbers measure different things. An assertion can check one sentinel byte, while an integration case may involve a complete process, load, fetch, and digest. Adding the numbers together would create a misleading impression of thousands of independent end-to-end tests.
The large assertion count mainly reflects byte-by-byte guard checks across the output-capacity loop. It is reported for reproducibility, not as a security score. A smaller number of carefully chosen failure cases can be more informative than many repeated assertions about the same behavior. The test design is therefore explained alongside the counts.
The harness launches sixteen independent digest processes using four worker threads in the Python controller. Each process has its own address space and provider instance. All results match their reference values. This demonstrates that the installed module and process-level invocation can be used concurrently without conflict in the tested scenario.
It does not test multiple threads sharing a provider context inside one application. It also does not test concurrent use of one digest context, which would be a separate state-sharing question. The machine-readable result explicitly records that shared_provider_context is false. This annotation prevents a later reader from describing the result as an in-process thread-safety test.
Each test stops on an unexpected result. Negative tests define success as the expected rejection rather than as a zero process exit. The harness does not silently continue and report a partial pass. A complete successful run writes a JSON result file and a final summary. The saved build log supplies the execution record used by the paper.
The test suite is intentionally deterministic and bounded. It is not a fuzzing campaign and has no probabilistic coverage claim. Future fuzzing should target a clearly defined input surface, such as parameter combinations or state transitions, and report its duration, corpus, instrumentation, and discovered failures separately. Reusing the word “tested” without those distinctions would hide important differences in assurance.