# Benchmark and performance evidence

Benchmark is controlled evidence about measured behavior. It is not a lifecycle
stage, a provider-selection mechanism, or a correctness proof.

## Controlled evidence, not a stage

Benchmark owns suite, case, workload, metric, sampling, environment, sample,
run, comparison, assessment, storage, and verification contracts. It owns the
bounded deterministic algorithms that derive summaries and confidence bounds,
the runner protocol and its conformance, and removable run-store providers.

The subject's contract or product owner retains workload meaning and correctness
laws. Benchmark does not own conformance, artifacts, Build-action identity,
process or remote execution, authority, provider admission, Planning selection,
generic telemetry, History facts, or product UI. A benchmark score is never
accessibility conformance, security proof, or semantic compatibility.

## Author one exact suite selection

Benchmark authoring uses the `one.test@1` domain's `benchmark` declaration.
`suite` is required, `case` is repeated, `policy <name> <ref>` names an exact
assessment policy, and a subject is either `subject candidate root <ref>` or
`subject <name> external <ref>` with an optional `optional` marker.

Illustrative syntax:

```one
one 1

semantic acme.orders.benchmark
    domain one.test@1
    benchmark OrderPlacement@1
        suite acme.orders.benchmarking#Placement@1
        policy qualification acme.orders.benchmarking#Baseline@1
        policy release acme.orders.benchmarking#ReleasePerformance@1
        subject candidate root acme.orders#OrdersApi@1
        subject legacy external acme.comparison#LegacyApi@1 optional
```

A comparison-only external subject never becomes a candidate merely because it
was measured. Its `optional` marker means absence is reported as unavailable,
not silently omitted or simulated.

The suite selection composes with an ordinary `one.system@1` root/component and
a `one.build@1` harness application that supplies the controlled executable:

```one
one 1

semantic acme.orders
    domain one.system@1
    component OrderPlacementHarness@1
        owner acme.orders#BenchmarkAuthors@1
        provide service acme.orders.benchmarking#Placement@1

    system AcmeOrders@1
        root orders_benchmark
            include OrderPlacementHarness
            build development
```

```one
one 1

semantic acme.orders.benchmark.harness
    domain one.build@1
    application OrderPlacementHarness@1
        profile <BenchmarkHarnessProfile>
        source ./bench/**
```

The `benchmark` declaration is exercised by the `one.test@1` normalizer fixture
(`fixture.performance#Startup@1`). The `acme.*` suite, root, and harness above
are illustrative end-product syntax.

## Environment and methodology are exact

`BenchmarkEnvironmentProfile` carries the validated, sealed environment that a
run claims: target, runtime and framework versions where applicable,
configuration, toolchain, and the exact bounds under which samples were taken.
It is validated, sealed, canonicalized, and decoded through the owning
contracts; an environment label never upgrades evidence.

`BenchmarkSamplingPolicy`, `BenchmarkExecutionOrder`, and `SamplePhase` define
sampling, ordered execution, warm-ups, and measurements. Raw samples preserve
acquisition order and are bounded. Aggregate calculation may sort a copy of the
evidence, but it never edits or filters the retained raw samples.

Removable `environment-local` and `environment-process` providers realize
execution and observation. They do not own workload meaning, select a provider,
or upgrade observed development evidence.

## Metrics and deterministic statistics

`BenchmarkMetric`, `BenchmarkAggregate`, `BetterDirection`, `BenchmarkUnit`,
and `MetricUsage` define what a value means and which direction is better.
Runner statistics are deterministic and named, including `summarize`,
`aggregate`, `confidence_interval`, `assess_absolute`, `assess_regression`,
`ensure_comparable`, and `NotComparableReason`.

Receipts keep each stage distinct: `BenchmarkSummary`, `BenchmarkRunReceipt`,
`BenchmarkAssessment`, `BenchmarkComparison`, and
`BenchmarkVerificationReceipt`. Missing, noisy, stale, incomparable, or
unverified evidence remains an exact disposition; it is never guessed or
silently omitted.

A regression assessment compares two runs of the same metric. For a metric
where smaller is better, the relative change is

$$
\Delta = \frac{\bar{x}_{\text{candidate}} - \bar{x}_{\text{baseline}}}{\bar{x}_{\text{baseline}}}
$$

and a regression gate accepts only $\Delta \le \tau$ for the declared tolerance
$\tau$. A negative $\Delta$ is an improvement. The point estimate alone never
decides: the confidence interval and comparability must also hold.

## Assessment policy and budgets

`BenchmarkAssessmentPolicy` selects absolute budgets (`AbsoluteBudget`) and
regression gates (`RegressionBudget`). `BenchmarkBudgetOutcome` records whether
a budget passed, failed, or could not be assessed; `AssessmentOutcome` records
the overall result.

Budgets are typed quantities. They compose with the cost model described in
[Cost and budgets](/one/platform/cost), and a benchmark result is only evidence
for the exact budget, workload, and environment it measured.

## Retention and independent verification

Runs are retained through a filesystem run store with explicit retain and load
operations. Retention is atomic and idempotent, follows no symlinks, detects
conflicts, and never exposes a partial run. Independent verification consumes
retained evidence and emits a `BenchmarkVerificationReceipt`.

## Planning consumes prior governed evidence

Hard semantic compatibility and required conformance are decided before any
performance or cost scoring. A benchmark result is evidence only for its exact
provider/profile and claim revisions, target and artifact revision, build
profile, framework/runtime/browser and versions where applicable, workload,
harness, configuration, and observation interval.

A benchmark can never repair semantic incompatibility or become a
candidate-capable realization. External comparison subjects never become
candidates. Planning may compare evidence only when its scope covers the
requested target and workload; missing required coverage is `MissingEvidence`.
A run made while realizing a plan cannot influence that same immutable plan.

See [Providers and adapters](/one/platform/providers) and
[Check, test, and invoke](/one/lifecycle/check-test-invoke).

## Bounds

- at most 256 cases per suite;
- at most 32 subjects and 32 metrics per case;
- at most 100 warm-ups and 4,096 measurements per series;
- at most 1,048,576 sample points per run;
- raw canonical sample data at most 64 MiB; every other canonical Benchmark
  document at most 16 MiB;
- case deadline at most one hour; suite deadline at most six hours;
- strings at most 1 KiB; stable diagnostics at most 4 KiB.

## Run a controlled suite

```console
one check --root systems/developer-interface/benchmarks/shell-startup
one test benchmark \
  --root systems/developer-interface/benchmarks/shell-startup \
  --policy qualification
```

The checked-in `shell-startup` root is a controlled harness whose numbers remain
illustrative until an owning benchmark run records them. Accessibility
conformance is never inferred from a benchmark score; follow
[Qualify a release](/one/examples/release-qualification) and
[tests, simulations, and fixtures](/one/authoring/testing) for those
obligations.

## Normative owner & evidence

The canonical owner is the
[Benchmark system](https://github.com/muijf/one/blob/main/systems/benchmark/AGENTS.md).
Planning's consumption boundary is owned by
[Planning](https://github.com/muijf/one/blob/main/systems/planning/AGENTS.md).

Demonstrated today: the `one.test@1` domain and its benchmark normalizer
fixture, the `one-benchmark` bounded atoms, the runner, the environment
providers, and the filesystem run-store and conformance suites. Illustrative:
the `acme.orders` suite/root/harness syntax, the console invocation above, and
any measurement not recorded by an owning benchmark run.
