18. Performance: A Proposed Evaluation Methodology

This chapter proposes how the design's performance would be evaluated. It reports no measurements, because none were taken. The distinction is maintained deliberately: a methodology is a contribution, and invented numbers are not.

18.1 Two costs, commonly conflated

Provider performance discussions usually merge two distinct costs.

Fetch cost is the price of resolving an algorithm name and property query to an implementation. The OpenSSL project states that this involves searching by name and is much slower than directly accessing a method table, recommending prefetching for algorithms used many times. It is paid per fetch.

Operation cost is the price of the cryptography itself, paid per byte or per operation.

For TLS these have entirely different profiles. A handshake performs a bounded number of fetches; a bulk transfer performs one AEAD operation per record for the life of the connection. Fetch cost therefore affects connection establishment rate, and operation cost affects throughput. A benchmark that does not separate them measures a mixture whose proportions depend on the connection lifetime it happened to choose.

18.2 Proposed measurements

#MeasuresMethodExpected sensitivity
M1Fetch latencyTime a single fetch, cold and warm, for each algorithmProperty query complexity; number of activated providers
M2Handshake rateComplete handshakes per second, one connection at a timeFetch cost; key exchange cost
M3ThroughputBytes per second over an established connectionAEAD operation cost only
M4Record-size sensitivityM3 across record sizes from small to maximumPer-call overhead relative to per-byte cost
M5Prefetch benefitM2 with and without prefetched algorithm objectsIsolates the cost N1 exists to avoid

M4 deserves comment. Per-call overhead is amortised over the record for large records and dominates for small ones, so a component with high fixed cost looks acceptable at maximum record size and poor for interactive traffic. Measuring one record size and reporting it as throughput conceals exactly the case that matters for latency-sensitive applications.

18.3 Baselines

Every measurement requires a comparison, and the meaningful baseline is the default provider's implementation of the nearest standard algorithm, measured on the same machine in the same session. Absolute figures are uninformative: they describe the test machine.

Where the design delegates (§7.4), a second baseline is valuable — the delegate measured directly, without the provider indirection. The difference between the two isolates the cost of the provider layer itself, which is the quantity a study of integration mechanics actually wants to know.

18.4 Controls

A measurement of this kind is easy to perform and hard to perform meaningfully. The proposed controls are:

18.5 Threats to validity

Stated in advance, as they should be for any proposed evaluation:

Implementation quality confound
A comparison against an accelerated standard implementation measures optimisation effort, not architecture. Any claim about the provider mechanism's overhead must isolate it as in §18.3.
Single-machine results
Cache sizes and instruction sets differ; results transfer poorly.
Loopback measurement
Handshake rate over loopback omits network latency, which dominates in reality. The figure is useful for comparing implementations and misleading as a prediction of deployed behaviour.
Microbenchmark bias
M1 in a tight loop measures a warm cache state that a real application will not have.

18.6 What would be reported

An executed study following this methodology would report, for each measurement, the distribution across runs, the baseline, the ratio with a confidence interval, and the environment. It would not report a single speed figure, and it would not compare against published numbers from other machines.

Until such a study is performed, the correct statement about this design's performance is that it is unmeasured. This chapter exists so that the measurement, when made, is made properly.