Skip to content
DocumentationBenchmark & performance evidence
On this page

Page resources

Open Markdownllms.txtView source

Last updated

Benchmark is controlled evidence about measured behavior. It is not a lifecycle stage, a provider-selection mechanism, or a correctness proof.

Controlled evidence, not a stage

Benchmark owns suite, case, workload, metric, sampling, environment, sample, run, comparison, assessment, storage, and verification contracts. It owns the bounded deterministic algorithms that derive summaries and confidence bounds, the runner protocol and its conformance, and removable run-store providers.

The subject's contract or product owner retains workload meaning and correctness laws. Benchmark does not own conformance, artifacts, Build-action identity, process or remote execution, authority, provider admission, Planning selection, generic telemetry, History facts, or product UI. A benchmark score is never accessibility conformance, security proof, or semantic compatibility.

Author one exact suite selection

Benchmark authoring uses the one.test@1 domain's benchmark declaration. suite is required, case is repeated, policy <name> <ref> names an exact assessment policy, and a subject is either subject candidate root <ref> or subject <name> external <ref> with an optional optional marker.

Illustrative syntax:

One
one 1

semantic acme.orders.benchmark
    domain one.test@1
    benchmark OrderPlacement@1
        suite acme.orders.benchmarking#Placement@1
        policy qualification acme.orders.benchmarking#Baseline@1
        policy release acme.orders.benchmarking#ReleasePerformance@1
        subject candidate root acme.orders#OrdersApi@1
        subject legacy external acme.comparison#LegacyApi@1 optional

A comparison-only external subject never becomes a candidate merely because it was measured. Its optional marker means absence is reported as unavailable, not silently omitted or simulated.

The suite selection composes with an ordinary one.system@1 root/component and a one.build@1 harness application that supplies the controlled executable:

One
one 1

semantic acme.orders
    domain one.system@1
    component OrderPlacementHarness@1
        owner acme.orders#BenchmarkAuthors@1
        provide service acme.orders.benchmarking#Placement@1

    system AcmeOrders@1
        root orders_benchmark
            include OrderPlacementHarness
            build development
One
one 1

semantic acme.orders.benchmark.harness
    domain one.build@1
    application OrderPlacementHarness@1
        profile <BenchmarkHarnessProfile>
        source ./bench/**

The benchmark declaration is exercised by the one.test@1 normalizer fixture (fixture.performance#Startup@1). The acme.* suite, root, and harness above are illustrative end-product syntax.

Environment and methodology are exact

BenchmarkEnvironmentProfile carries the validated, sealed environment that a run claims: target, runtime and framework versions where applicable, configuration, toolchain, and the exact bounds under which samples were taken. It is validated, sealed, canonicalized, and decoded through the owning contracts; an environment label never upgrades evidence.

BenchmarkSamplingPolicy, BenchmarkExecutionOrder, and SamplePhase define sampling, ordered execution, warm-ups, and measurements. Raw samples preserve acquisition order and are bounded. Aggregate calculation may sort a copy of the evidence, but it never edits or filters the retained raw samples.

Removable environment-local and environment-process providers realize execution and observation. They do not own workload meaning, select a provider, or upgrade observed development evidence.

Metrics and deterministic statistics

BenchmarkMetric, BenchmarkAggregate, BetterDirection, BenchmarkUnit, and MetricUsage define what a value means and which direction is better. Runner statistics are deterministic and named, including summarize, aggregate, confidence_interval, assess_absolute, assess_regression, ensure_comparable, and NotComparableReason.

Receipts keep each stage distinct: BenchmarkSummary, BenchmarkRunReceipt, BenchmarkAssessment, BenchmarkComparison, and BenchmarkVerificationReceipt. Missing, noisy, stale, incomparable, or unverified evidence remains an exact disposition; it is never guessed or silently omitted.

A regression assessment compares two runs of the same metric. For a metric where smaller is better, the relative change is

Δ=xˉcandidate−xˉbaselinexˉbaseline\Delta = \frac{\bar{x}_{\text{candidate}} - \bar{x}_{\text{baseline}}}{\bar{x}_{\text{baseline}}}

and a regression gate accepts only Δ≤τ\Delta \le \tau for the declared tolerance τ\tau. A negative Δ\Delta is an improvement. The point estimate alone never decides: the confidence interval and comparability must also hold.

Assessment policy and budgets

BenchmarkAssessmentPolicy selects absolute budgets (AbsoluteBudget) and regression gates (RegressionBudget). BenchmarkBudgetOutcome records whether a budget passed, failed, or could not be assessed; AssessmentOutcome records the overall result.

Budgets are typed quantities. They compose with the cost model described in Cost and budgets, and a benchmark result is only evidence for the exact budget, workload, and environment it measured.

Retention and independent verification

Runs are retained through a filesystem run store with explicit retain and load operations. Retention is atomic and idempotent, follows no symlinks, detects conflicts, and never exposes a partial run. Independent verification consumes retained evidence and emits a BenchmarkVerificationReceipt.

Planning consumes prior governed evidence

Hard semantic compatibility and required conformance are decided before any performance or cost scoring. A benchmark result is evidence only for its exact provider/profile and claim revisions, target and artifact revision, build profile, framework/runtime/browser and versions where applicable, workload, harness, configuration, and observation interval.

A benchmark can never repair semantic incompatibility or become a candidate-capable realization. External comparison subjects never become candidates. Planning may compare evidence only when its scope covers the requested target and workload; missing required coverage is MissingEvidence. A run made while realizing a plan cannot influence that same immutable plan.

See Providers and adapters and Check, test, and invoke.

Bounds

  • at most 256 cases per suite;
  • at most 32 subjects and 32 metrics per case;
  • at most 100 warm-ups and 4,096 measurements per series;
  • at most 1,048,576 sample points per run;
  • raw canonical sample data at most 64 MiB; every other canonical Benchmark document at most 16 MiB;
  • case deadline at most one hour; suite deadline at most six hours;
  • strings at most 1 KiB; stable diagnostics at most 4 KiB.

Run a controlled suite

Console
one check --root systems/developer-interface/benchmarks/shell-startup
one test benchmark \
  --root systems/developer-interface/benchmarks/shell-startup \
  --policy qualification

The checked-in shell-startup root is a controlled harness whose numbers remain illustrative until an owning benchmark run records them. Accessibility conformance is never inferred from a benchmark score; follow Qualify a release and tests, simulations, and fixtures for those obligations.

Normative owner & evidence

The canonical owner is the Benchmark system. Planning's consumption boundary is owned by Planning.

Demonstrated today: the one.test@1 domain and its benchmark normalizer fixture, the one-benchmark bounded atoms, the runner, the environment providers, and the filesystem run-store and conformance suites. Illustrative: the acme.orders suite/root/harness syntax, the console invocation above, and any measurement not recorded by an owning benchmark run.