This chapter proposes how the design's performance would be evaluated. It reports no measurements, because none were taken. The distinction is maintained deliberately: a methodology is a contribution, and invented numbers are not.
Provider performance discussions usually merge two distinct costs.
Fetch cost is the price of resolving an algorithm name and property query to an implementation. The OpenSSL project states that this involves searching by name and is much slower than directly accessing a method table, recommending prefetching for algorithms used many times. It is paid per fetch.
Operation cost is the price of the cryptography itself, paid per byte or per operation.
For TLS these have entirely different profiles. A handshake performs a bounded number of fetches; a bulk transfer performs one AEAD operation per record for the life of the connection. Fetch cost therefore affects connection establishment rate, and operation cost affects throughput. A benchmark that does not separate them measures a mixture whose proportions depend on the connection lifetime it happened to choose.
| # | Measures | Method | Expected sensitivity |
|---|---|---|---|
| M1 | Fetch latency | Time a single fetch, cold and warm, for each algorithm | Property query complexity; number of activated providers |
| M2 | Handshake rate | Complete handshakes per second, one connection at a time | Fetch cost; key exchange cost |
| M3 | Throughput | Bytes per second over an established connection | AEAD operation cost only |
| M4 | Record-size sensitivity | M3 across record sizes from small to maximum | Per-call overhead relative to per-byte cost |
| M5 | Prefetch benefit | M2 with and without prefetched algorithm objects | Isolates the cost N1 exists to avoid |
M4 deserves comment. Per-call overhead is amortised over the record for large records and dominates for small ones, so a component with high fixed cost looks acceptable at maximum record size and poor for interactive traffic. Measuring one record size and reporting it as throughput conceals exactly the case that matters for latency-sensitive applications.
Every measurement requires a comparison, and the meaningful baseline is the default provider's implementation of the nearest standard algorithm, measured on the same machine in the same session. Absolute figures are uninformative: they describe the test machine.
Where the design delegates (§7.4), a second baseline is valuable — the delegate measured directly, without the provider indirection. The difference between the two isolates the cost of the provider layer itself, which is the quantity a study of integration mechanics actually wants to know.
A measurement of this kind is easy to perform and hard to perform meaningfully. The proposed controls are:
Stated in advance, as they should be for any proposed evaluation:
An executed study following this methodology would report, for each measurement, the distribution across runs, the baseline, the ratio with a confidence interval, and the environment. It would not report a single speed figure, and it would not compare against published numbers from other machines.
Until such a study is performed, the correct statement about this design's performance is that it is unmeasured. This chapter exists so that the measurement, when made, is made properly.