8. Expanded Research Design

The baseline experiment answers whether the teaching provider can be made to work through several ordinary interfaces. A longer engineering study must ask a more demanding question: what exactly does “works” mean, and how should evidence be organized so that another reader can evaluate the claim? This chapter develops a requirements model and an evidence model for the extended investigation. The additional work does not turn the bridge into a new cryptographic primitive. It increases the precision with which its integration behavior is described.

8.1 Unit of analysis

The unit of analysis is the combination of the module, the host OpenSSL library, the configuration used by a particular process, and the caller’s selection request. Examining the shared object alone would miss important interactions. For example, a valid module can remain unavailable because the application reads a different configuration file. Conversely, a correct digest can be returned by another provider even when the intended module was never selected. The experiment therefore treats observable application behavior as an outcome of several cooperating components.

This choice also determines the meaning of reproducibility. Copying the C source is necessary but insufficient. A reproducing investigator needs the compiler invocation, the development headers, the runtime version, the module path, and the environment used for the test. The saved artifacts supply these details or provide commands that regenerate them. Where the host still contributes a dependency, such as the system compiler, the paper reports that dependency rather than presenting the directory as a complete operating-system image.

8.2 Functional and nonfunctional requirements

Functional requirements describe responses that can be observed directly. The provider should load, advertise its digest, produce the expected bytes, accept streaming input, and reject an unsatisfied selection constraint. Nonfunctional requirements concern qualities such as maintainability, diagnosability, compatibility, and confidence in resource management. They often require more than a single successful test. A readable ownership table is evidence of design discipline, but it is not equivalent to a memory-safety proof.

IDRequirementEvidence used
FR1Expose EDU-SHA256 through the provider interface.Programmatic fetch and algorithm-table inspection.
FR2Return 32 correct output bytes for tested messages.Reference comparisons and direct finalization checks.
FR3Support streaming and operation-state duplication.C client that branches a partially updated context.
FR4Respect outer mandatory property matching.Positive and negative fetch requests.
FR5Load from an explicit configuration file.Controlled process environment and activation matrix.
NFR1Make ownership reviewable.Design tables and cleanup-path inspection.
NFR2Make experiments reproducible.Scripts, source snapshots, and machine-readable results.
NFR3Avoid unsupported assurance claims.Explicit separation of observations and limitations.

The matrix deliberately does not label NFR1 “proved.” The direct tests exercise many normal allocations and releases, but they do not force every allocator to fail at every call. Similarly, NFR2 is supported on the recorded host; it is not evidence that every operating system can compile the module unchanged. This distinction is useful in undergraduate engineering because it prevents test completion from being mistaken for complete requirement verification.

8.3 Competing explanations

A digest match has several possible explanations. The module may have executed correctly; another provider may have executed instead; the input may have been altered in the harness; or the reference may share the same defect as the implementation. The design reduces some of these ambiguities. A distinct algorithm name and an explicit provider property reduce accidental selection. Binary input and output avoid newline and hexadecimal-format assumptions. The client checks the provider attached to the fetched method. None of these steps makes the backend independent of OpenSSL, so correlated algorithm errors remain a limitation.

A failed command is similarly ambiguous. Failure could be the desired property rejection or an unrelated loading problem. The negative property case is run beside a successful case with the same executable, module path, and message. Only the property changes. The difference supports the intended explanation, although a richer error-classification harness could make the inference stronger. The paper reports the measured exit code rather than claiming to have exhaustively classified every OpenSSL error-stack entry.

8.4 Evidence levels

Four evidence levels are distinguished throughout the extended discussion. Documentary evidence describes the public interface as specified in a version-pinned manual. Inspection evidence follows the prototype’s control flow and resource ownership. Execution evidence records what a concrete test actually did. Proposed evidence describes tests or operational reviews that would be appropriate in future work but were not performed. These levels should not be silently substituted for one another.

For example, the baseline paper inspected the finalization guard but did not invoke it with an undersized buffer. The extension adds execution evidence for capacities from zero through 33 bytes. This is a real increase in coverage, not a wording change. Yet the same extension still does not provide allocator-failure injection. The appropriate conclusion is specific: the tested buffer guards behaved as intended; exhaustive failure-path assurance remains open.

8.5 Scope of the extended study

The extension retains the original provider code so that additional evidence can be attributed to a fixed implementation. It adds a white-box callback harness and a broader integration harness. The white-box harness directly invokes the prototype’s entry point with null core arguments because this implementation does not consume them. That is a controlled property of this prototype, not a general technique guaranteed to work for arbitrary providers. The integration harness continues to use ordinary OpenSSL command-line requests.

The resulting study therefore has two complementary perspectives. One observes the module as an application would use it. The other deliberately reaches the provider-facing interface to test conditions that an ordinary EVP call may filter or normalize before they reach the module. Agreement between these perspectives is useful, but neither eliminates the need to understand the specific layer at which an assertion is made.