Benchmark is controlled evidence about measured behavior. It is not a lifecycle stage, a provider-selection mechanism, or a correctness proof.
Controlled evidence, not a stage
Benchmark owns suite, case, workload, metric, sampling, environment, sample, run, comparison, assessment, storage, and verification contracts. It owns the bounded deterministic algorithms that derive summaries and confidence bounds, the runner protocol and its conformance, and removable run-store providers.
The subject's contract or product owner retains workload meaning and correctness laws. Benchmark does not own conformance, artifacts, Build-action identity, process or remote execution, authority, provider admission, Planning selection, generic telemetry, History facts, or product UI. A benchmark score is never accessibility conformance, security proof, or semantic compatibility.
Author one exact suite selection
Benchmark authoring uses the one.test@1 domain's benchmark declaration.
suite is required, case is repeated, policy <name> <ref> names an exact
assessment policy, and a subject is either subject candidate root <ref> or
subject <name> external <ref> with an optional optional marker.
Illustrative syntax:
one 1
semantic acme.orders.benchmark
domain one.test@1
benchmark OrderPlacement@1
suite acme.orders.benchmarking#Placement@1
policy qualification acme.orders.benchmarking#Baseline@1
policy release acme.orders.benchmarking#ReleasePerformance@1
subject candidate root acme.orders#OrdersApi@1
subject legacy external acme.comparison#LegacyApi@1 optionalA comparison-only external subject never becomes a candidate merely because it
was measured. Its optional marker means absence is reported as unavailable,
not silently omitted or simulated.
The suite selection composes with an ordinary one.system@1 root/component and
a one.build@1 harness application that supplies the controlled executable:
one 1
semantic acme.orders
domain one.system@1
component OrderPlacementHarness@1
owner acme.orders#BenchmarkAuthors@1
provide service acme.orders.benchmarking#Placement@1
system AcmeOrders@1
root orders_benchmark
include OrderPlacementHarness
build developmentone 1
semantic acme.orders.benchmark.harness
domain one.build@1
application OrderPlacementHarness@1
profile <BenchmarkHarnessProfile>
source ./bench/**The benchmark declaration is exercised by the one.test@1 normalizer fixture
(fixture.performance#Startup@1). The acme.* suite, root, and harness above
are illustrative end-product syntax.
Environment and methodology are exact
BenchmarkEnvironmentProfile carries the validated, sealed environment that a
run claims: target, runtime and framework versions where applicable,
configuration, toolchain, and the exact bounds under which samples were taken.
It is validated, sealed, canonicalized, and decoded through the owning
contracts; an environment label never upgrades evidence.
BenchmarkSamplingPolicy, BenchmarkExecutionOrder, and SamplePhase define
sampling, ordered execution, warm-ups, and measurements. Raw samples preserve
acquisition order and are bounded. Aggregate calculation may sort a copy of the
evidence, but it never edits or filters the retained raw samples.
Removable environment-local and environment-process providers realize
execution and observation. They do not own workload meaning, select a provider,
or upgrade observed development evidence.
Metrics and deterministic statistics
BenchmarkMetric, BenchmarkAggregate, BetterDirection, BenchmarkUnit,
and MetricUsage define what a value means and which direction is better.
Runner statistics are deterministic and named, including summarize,
aggregate, confidence_interval, assess_absolute, assess_regression,
ensure_comparable, and NotComparableReason.
Receipts keep each stage distinct: BenchmarkSummary, BenchmarkRunReceipt,
BenchmarkAssessment, BenchmarkComparison, and
BenchmarkVerificationReceipt. Missing, noisy, stale, incomparable, or
unverified evidence remains an exact disposition; it is never guessed or
silently omitted.
A regression assessment compares two runs of the same metric. For a metric where smaller is better, the relative change is
and a regression gate accepts only for the declared tolerance . A negative is an improvement. The point estimate alone never decides: the confidence interval and comparability must also hold.
Assessment policy and budgets
BenchmarkAssessmentPolicy selects absolute budgets (AbsoluteBudget) and
regression gates (RegressionBudget). BenchmarkBudgetOutcome records whether
a budget passed, failed, or could not be assessed; AssessmentOutcome records
the overall result.
Budgets are typed quantities. They compose with the cost model described in Cost and budgets, and a benchmark result is only evidence for the exact budget, workload, and environment it measured.
Retention and independent verification
Runs are retained through a filesystem run store with explicit retain and load
operations. Retention is atomic and idempotent, follows no symlinks, detects
conflicts, and never exposes a partial run. Independent verification consumes
retained evidence and emits a BenchmarkVerificationReceipt.
Planning consumes prior governed evidence
Hard semantic compatibility and required conformance are decided before any performance or cost scoring. A benchmark result is evidence only for its exact provider/profile and claim revisions, target and artifact revision, build profile, framework/runtime/browser and versions where applicable, workload, harness, configuration, and observation interval.
A benchmark can never repair semantic incompatibility or become a
candidate-capable realization. External comparison subjects never become
candidates. Planning may compare evidence only when its scope covers the
requested target and workload; missing required coverage is MissingEvidence.
A run made while realizing a plan cannot influence that same immutable plan.
See Providers and adapters and Check, test, and invoke.
Bounds
- at most 256 cases per suite;
- at most 32 subjects and 32 metrics per case;
- at most 100 warm-ups and 4,096 measurements per series;
- at most 1,048,576 sample points per run;
- raw canonical sample data at most 64 MiB; every other canonical Benchmark document at most 16 MiB;
- case deadline at most one hour; suite deadline at most six hours;
- strings at most 1 KiB; stable diagnostics at most 4 KiB.
Run a controlled suite
one check --root systems/developer-interface/benchmarks/shell-startup
one test benchmark \
--root systems/developer-interface/benchmarks/shell-startup \
--policy qualificationThe checked-in shell-startup root is a controlled harness whose numbers remain
illustrative until an owning benchmark run records them. Accessibility
conformance is never inferred from a benchmark score; follow
Qualify a release and
tests, simulations, and fixtures for those
obligations.
Normative owner & evidence
The canonical owner is the Benchmark system. Planning's consumption boundary is owned by Planning.
Demonstrated today: the one.test@1 domain and its benchmark normalizer
fixture, the one-benchmark bounded atoms, the runner, the environment
providers, and the filesystem run-store and conformance suites. Illustrative:
the acme.orders suite/root/harness syntax, the console invocation above, and
any measurement not recorded by an owning benchmark run.