Back to explorer

We measure the tool, not the agent. Most agent benchmarks hold the task fixed and vary the model. This benchmark inverts that: one constant reference agent, the same task per category, five trials — so a difference in the numbers is attributable to the tool and how the agent was taught it, not the setup.

Agent Ergonomics — Measurement Methodology v0

Status: frozen · Version: v0 · the hash of this file is recorded in docs/methodology.lock and stamped into every profile's provenance.methodology. A profile whose hash does not match the current file was produced under a different methodology and is not comparable — regenerate it.

Why this document exists: a benchmark is only a benchmark if it is reproducible and its rules are disclosed before the numbers. This is the Artificial-Analysis discipline applied to Agent Experience: fix the harness, vary the subject, prefer ground-truth signals to judgement, keep axes separate, report variance.


1. What is measured

The subject is a tool/skill/MCP-service an agent uses to produce an artifact. The subject is the ONLY thing that varies between profiles. Everything else — the agent, the task difficulty, the trial count, the scoring — is held constant so a difference in the numbers is attributable to the tool, not the setup.

Output is the surfaces × lenses matrix (5 × 6 = 30 cells) plus roll-ups. See ax-model.md for the model. A cell is emitted only when it was actually measured; unmeasured cells are reported as gaps, never filled from nothing.

2. The reference agent (the constant)

  • Reference agent: the PI coding agent (@earendil-works/pi-coding-agent), model z-ai/glm-5.2 over OpenRouter, run non-interactively (pi --print --mode json). One agent scores every profile.
  • Rationale: it reports real token telemetry (input/output/cost per turn), which anchors the economy lens — the one lens with the widest real spread.
  • The reference agent is the constant. Measuring "does AX hold across different agents?" is a separate study (ax deep … --ensemble), whose cross-model runs never feed the scored matrix.

3. Subjects & provisioning contract

  • A subject key is <category>-<tool> (e.g. diagrams-d2). No oota: prefix.
  • Resolution order: provisioned skill (skills.sh / official, via npx skills add) → provisioned MCP server → the tool's own source.
  • Provisioning must succeed for the mode it claims. A subject that silently falls back to source while its peers ran from an installed skill is not comparable and must be flagged, not scored alongside them. (provisioned.json records the mode; profiles record skillMode.)

4. Trials

  • N = 5 trials per tier (--n overrides; never below 2 for any tier whose reading needs variance).
  • Per-trial isolation: each trial runs in a fresh copy of the grounded cwd. Trials never share mutable state, so run-to-run determinism is a real measurement rather than an artifact of one directory being mutated in place.
  • Every reading carries std across the N trials. Variance is reported, not hidden.

5. Grounded briefs

A brief = (category task-spec) × (a real seed entity supplying the subject matter). Same category → same seed by default (controlled comparison). The task-spec is checkable (see §7). Briefs are generated deterministically (no LLM), so the bar is identical for every tool in a category.

6. The probe battery (what each tier ACTUALLY does)

Each tier is a prompt sequence (± a cwd mutation), not the same one-shot intent relabelled. Prompts run over a persistent cwd (PI: sequential invocations, state persists on disk; ACP: one open session). intentTemplate is threaded; scenarios live in ax/probes/scenarios.ts.

Tier Interaction Executable when Cells (dropped if not executable)
T0 cold-call from memory, no docs, no tools always interface.prior_alignment
T1 guided one prompt, docs available always loop.verifiability/economy/determinism, interface.verifiability
T2 multi-turn base + 2 evolving follow-ups over persistent cwd always recursion.coherence (needs checker), recursion.economy (needs tokens)
T3 failure-injection a broken artifact is seeded; agent must recover it broken template exists for the tool loop.safety
T5 compaction context inflated + long iterative task only if real compaction occurs recursion.determinismdropped when compactionObserved = 0
T6 refinement base + "push past acceptable" ×2 needs checker recursion.verifiability

Rule: if a tier cannot execute what it claims against the reference agent, its cells are deleted, not filled from a run that didn't do the thing. (Concretely: T5 determinism requires observed compaction; goal-based cells require a checker; token-based cells require token telemetry.)

7. The ground-truth hierarchy (checker → telemetry → judge → rater)

Signals are trusted in this order; a lower tier only fills what a higher tier can't:

  1. Programmatic checker (ax/instruments/checker.ts) — deterministic, no LLM. Given the category + the produced artifact, it computes which of the brief's checkable requirements are met. Its score (fraction met, 0..1) is goalProgress — the trajectory signal AND the anchor the rater is calibrated against (never the reverse). Checkers wired (the CHECKERS registry):

    • diagrams — d2 · mermaid · plantuml · dot · python-diagrams (renders + node/ edge/cluster/title counts); excalidraw (JSON: shapes, bound arrows, frames, title). Live-validated at N≥5.
    • data-visualization — a figure was produced, its source runs clean, and (when SVG) it carries labels + multiple series colours. Built; pending live validation.
    • documents — compiles to final format (tectonic/pdflatex/pandoc/typst) + 3+ sections + table + figure + citation. Built; pending live validation.
    • markup — renders + title + nested sections + table + code block + link + image. Built; pending live validation.
    • image-processing — WebP + JPEG outputs, cropped ~1.91:1, each <200KB (dims via sips/identify, size on disk). Built; pending live validation.

    A category is v0-validated once it has been run live at N≥5 with clean provisioning; until then its checker is active but the category's numbers are provisional. Whole-file hash of THIS doc pins the checker inventory, so adding a checker is a deliberate, versioned change.

  2. Telemetry — tokens, turns, success, compaction, run-to-run variance. Real, parsed from the agent's event stream. No estimation.

  3. Artifact judge (judge.ts) — a modality-aware vision panel that SEES the artifact, used where a programmatic checker doesn't exist. Secondary.

  4. Rater panel (rater.ts) — LLM judges fill cells with no hard instrument, from the real transcript/docs/telemetry. Independence: the panel MUST NOT include the reference model (a model grading its own work is not independent); the reference z-ai/glm-5.2 is excluded. The panel is handed the checker's verdict as ground truth and told to defer to it on verifiability/determinism.

A category with no checker returns hasChecker:false → its goal-based cells stay soft (judge/rater only), at reduced confidence, and are flagged.

8. Calibration & validity

  • Convergent validity: where a hard instrument and the independent rater score the same cell, the delta is recorded (profile.validation). Per-lens rater bias is measured and subtracted from that lens's soft cells (anchoring soft to hard).
  • Peer calibration: absolute −1..+1 scores read as "everything green" (scoring-audit.md). The published standing is percentile vs. the measured field, with low-variance ("dead") cells down-weighted. This sharpens as N subjects grows and requires like-vs-like peers to be trustworthy.
  • Construct validity: ax perturb degrades one dimension (e.g. strips docs) and checks the matrix moves in the right cells and stays put elsewhere.

9. Reproducibility & cost

  • Same subject + same methodology hash + same seed → same brief and same scoring procedure. Randomness in the agent is bounded by N and reported as std.
  • Transcripts + per-trial telemetry are persisted for re-analysis without re-running.
  • Cost controls: Screen (static + T0/T1) before Deep; cap N at 5 unless a comparison demands more.

10. Known limitations (honest)

  • Checkers exist only for diagrams so far. Every other category's goal-based cells are soft (judge/rater). Breadth without a checker multiplies rater noise — see §11.
  • Compaction (T5) is hard to force reproducibly with a one-shot reference agent; those cells are frequently (honestly) dropped.
  • Peer percentiles need N subjects and like-vs-like peers. At small N they are coarse. A field of weak tools makes a mediocre one look green; the checker + perturbation guards the absolute anchor.

11. Scaling gate (Part 5 — deferred on purpose)

Adding services/skills multiplies whatever the instrument is. Breadth is unlocked per category only after that category has: (a) a programmatic checker, (b) a provisioning path that never silently falls back to source (enforced — provisionFallback/comparable flags in provenance), (c) N ≥ 2 with reported variance. Diagrams clears this gate (live-validated).

Checkers now exist for data-visualization, documents, markup, image-processing (§7) — the (a) requirement is met for these. Remaining before their numbers are promoted from provisional to v0-validated: a clean provisioning entry in skills.curated.json for at least one tool per category, and a live N≥5 deep run. Only then does skills.curated.json expand further across the 29 categories, and MCP services beyond excalidraw get added. Order stays: checker → provisioning → live-validate → expand, never expand-first.


Changelog

v0.1 (2026-07-06) — measurement sharpened

Changes that alter default scoring, hence a new hash. Profiles under the prior hash (57cb8d0be70fb17a, the first diagram batch) are a valid v0 baseline but are NOT comparable to v0.1 numbers — regenerate to compare.

  • Determinism = pass^k (ax/instruments/reliability.ts) — the τ-bench estimator P(k random trials all pass), replacing the old success/turn-variance heuristic. Sharper, standard, checker-grounded where available.
  • Human surface instrumented (ax/instruments/human.ts) — human.safety (reversibility/oversight), human.verifiability (transparency), human.coherence (steerable control surface) are now scored from checkable patterns (Microsoft Agentic Design Principles) instead of resting entirely on the rater.
  • OTel GenAI adapter (ax/harness/otel.ts, docs/otel-telemetry-contract.md) — any OpenTelemetry-instrumented harness can feed the matrix; PI is one adapter.
  • Recursion.determinism via FRESH-CONTEXT CONTINUATION — the compaction-survival signal now comes from T2: each segment is a cold agent (no conversation history) seeing only the cwd, so if the tool's artifacts aren't self-describing, a cold agent breaks prior work and goalProgress regresses. Stability across segments = the score. This retires T5's compaction-forcing entirely (it cost ~43 min/tool for compaction=0 on a one-shot agent — pure waste). T5 remains defined/opt-in (--tiers T5) for genuinely session-continuous agents, but is out of the default.
  • Parallel trials — the N independent trials of a tier run CONCURRENTLY (OOTA_TRIAL_CONCURRENCY, default 5), not serially; PI switched from blocking spawnSync to async spawn. A tier's wall-time drops from N× to ≈ the slowest single trial. (The remaining floor is per-call agent latency — configurable cap OOTA_PI_TIMEOUT_MS, with timeouts logged, not silently absorbed.)
  • Merge preserves evidencecollapse() no longer discards raw/std/n when two instruments fill the same cell (pass^k curve, token counts, checker detail now survive).
  • OTel ingest is liveax otel <trace.otlp.json> <subject> [--cwd dir] scores telemetry cells from any OpenTelemetry GenAI trace (not just a bare adapter).
  • Incremental persistenceax deep writes a partial profile after every tier, so a long/killed run is never a total loss.
  • Task-suite quarry (docs/task-suites.md) — where a validated external benchmark has a programmatic verifier, adapt it (verifier→checker, tasks→briefs); reference only, gated per the rules above.

Teaching sources — the source is a variable of AX

A tool doesn't serve an agent in the abstract; it serves the agent through a teaching source — an official skill, an official MCP, a library-specific community skill, or the tool's official documentation. The same tool taught two different ways gives two different ergonomics. So the AX subject is a (tool × source) pair, and comparing sources is a first-class goal, not noise.

Source kinds (all must be tool-specific)

Ranked by authority, but all are legitimate as long as they are specific to that library — never generic ("a diagram skill") and never our own hand-rolled wrapper:

  1. official-skill — a skill published by the tool's vendor (e.g. Resend's react-email).
  2. official-mcp — the tool's official MCP server (e.g. the Excalidraw MCP).
  3. official-docs — the tool's own documentation, pulled into markdown (fixtures/official-docs/<tool>.md). The always-available baseline when no official skill/MCP exists — the official information is present, not absent.
  4. community-skill — a library-specific skill from skills.sh, used when no official skill exists. Must be for that library (e.g. d2-diagram-creator for D2), highest-adoption preferred, and its rationale recorded in skills.curated.json.

Explicitly excluded: generic/catch-all skills, and any hand-rolled wrapper of ours (the former OOTA wrappers). Those measure our implementation, not the tool.

Addressing a source

diagrams-d2          # curated primary (here: the tool-specific community skill)
diagrams-d2@docs     # the official D2 documentation
diagrams-d2@skill    # the tool-specific skill, explicitly
diagrams-d2@mcp      # the official MCP, explicitly

Each (tool × source) produces its own profile, keyed by the suffixed id (diagrams-d2@docs), so multiple sources for one tool coexist and rank independently on the leaderboard.

Comparing sources (the "over time" goal)

Because the harness, reference agent, briefs, and N are held constant, a difference between diagrams-d2@docs and diagrams-d2@skill is attributable to the source — i.e. which way of teaching the tool serves the agent better. Two uses:

  • Per-source ranking — "for D2, the community skill beats the raw docs on Disclosure (+0.4) but is even on Loop." Actionable: it tells you whether a skill is worth its maintenance, and where docs fall short.
  • Tool-level roll-up — a tool's AX can be summarized as the best available source (what a well-set-up agent would actually use) and/or the mean across sources (how the tool tends to land however it's taught). Report both; never silently average away a great skill or a broken one.

As more skills accrue for a tool over time, this becomes a longitudinal study: the spread across sources is itself a signal (a tool that only works well via one bespoke skill is more fragile than one that's ergonomic from its official docs).

Status

  • official-docs pulled for d2, mermaid, plantuml, matplotlib (the diagram + data-viz tools that have no official skill).
  • Resolution + mounting live (resolveSubject, <key>@docs).
  • Cross-source comparison instrument: designed here, not yet built — the data (per-source profiles) is what it consumes; wire it once ≥2 sources per tool have been run.

The workflow benchmark. The above measures a tool in isolation. The agency benchmark measures tools and skills as workers inside real multi-stage projects for one fictional client, where each worker's input is another's real output and the whole delivery is judged as one body of work. Its rules follow.

Agency Benchmark Methodology — v1.0

(The workflow benchmark. The per-tool matrix methodology lives in methodology-v0.md and remains frozen — its hash is pinned into matrix profiles. This document governs everything under ax/agency/ and fixtures/agency/; agency run manifests pin THIS document's hash.)

1. What is measured

Not "can tool X produce artifact Y in isolation" (the matrix answers that), but how tools and skills perform as workers inside a real multi-stage project — where their input is another worker's real output, their output feeds the next worker, and the delivery is judged as one body of work. The unit of measurement is the chain, and the subjects are the same (tool × teaching-source) workers the matrix measures, plus skill-subjects (the skill IS the worker) and @bare baselines (no teaching at all).

2. The world: one client, one source of truth

Every project serves one fictional client — Verdant Labs, maker of the Canopy smart hydroponic garden — defined entirely by fixtures/universe/verdant/world.jsonl: company, legal/IP, products with SKUs and prices, system architecture with modules and known debt, people, customers, quarterly financials, pod telemetry, brand system, a 2026 roadmap with REQ-ids and priorities, an incident postmortem, a B2B order, and user personas.

The world-file is the single source of truth: brief material is derived from it (ax/briefs/universe.ts), and the coherence checker verifies deliverables against it — a quoted revenue figure, an invoice subtotal, or a brand hex is either the canonical value or a fabrication. A hardware+SaaS client was chosen deliberately so electronics, CAD, audio, and photography tasks arise as naturally as documents and dashboards. (Domain bias from the single universe is a known limitation; a second client in a different industry is planned as a rotation.)

3. Tickets: deterministic, dynamic projects

A ticket is a project generated from (template × seed)makeTicket(template, seed) is a pure function, so every ticket replays identically and different seeds diverge. Six templates order the work differently (no frozen chain): launch (ship a roadmap REQ), incident (respond to the postmortem), audit (quarterly board pack), expansion (win a persona segment), and two deliberately short shapes, quote and status.

Stage inclusion is conditional on the task content, not the template: an architecture pass exists only when the picked feature spans ≥2 modules; a chart only when the ticket references metrics; customer comms only when the change is customer-visible; plus seeded optional stages. Realism without frozen workflows; reproducibility without hand-picking.

4. Stages: typed artifacts, rostered workers, forced rotation

Stages compose via typed artifact interfaces (plan-doc, diagram, chart, web-ui, deck, doc, social-image, business-doc, infographic) — a stage sees its predecessors' artifacts, never their implementation, which makes chains language-agnostic by construction. Code-level integration is out of scope.

Each stage type has a roster of eligible workers. Assignment is a seeded Latin square: run k of a ticket uses roster[(hash(ticket, stage) + k) mod n] — deterministic, and provably no worker monopolizes a stage. Which workers served where, under which model, is bookkept in the coverage manifest (fixtures/agency/coverage.json), not enforced by mutable quotas.

5. Execution: live chains

Stages run in order. Each stage's worker gets: the client's material facet, its own teaching source (docs / skill mount / nothing for @bare), the ticket mission, its stage spec, and ./inputs/ containing the real deliverables of the stages it consumes. Nothing is simulated; a failed stage doesn't halt the chain (downstream workers cope with missing inputs, as in a real project). Each stage gets one shot with a 360-second budget (AX_PI_TIMEOUT_MS; the matrix's 180s default starved one-shot build stages — measured, then fixed).

Drivers span models and harnesses: the PI baseline (glm reference, plus e.g. MiMo via pi:<model>) and real coding harnesses over ACP (Claude Code + Sonnet via acp:<cmd>:<model>). Cross-model comparisons hold runIndex constant, so the same workers serve the same stages — only the driving model/harness changes.

6. Scoring: hard anchors first, judges on top

  1. Per-stage ground truth — the same deterministic checkers as the matrix (render + structural requirements per artifact type; plan-docs get their own checker: sections, tables, P0-P2 priorities, numeric targets).
  2. Bundle coherence (deterministic) — greps and arithmetic against the world-file: brand palette propagated into visual deliverables, ticket REQ/incident ids traced through the bundle, invoice math vs canonical SKU prices, quoted financials matching the canonical figures, build echoing the plan's REQ-ids. Checks are shape-conditioned: a shape with no visual stage is not penalized for carrying no hexes; the ref-tracing bar scales with deliverable count. (The unconditioned version structurally punished short shapes — caught and retrofitted.)
  3. chainScore = meanStageChecker × (0.5 + 0.5 × coherence) — the run's hard number; a chain of individually-fine artifacts that contradict each other scores poorly, as it should.
  4. Frontier bundle judge (best-effort) — a panel (Claude Opus 4.8 + Gemini 3.5 Flash) reviews the whole delivery for accomplishment, cross-deliverable agreement, and usability. Independence rule: a run's own driver model is excluded from its panel. Judges that error are skipped, never faked; with zero judges the hard scores stand alone.

7. Position fairness: depth is a covariate, not a confound

End-anchored stage types (billing) are structurally deeper. Every stage result records depth (# prior stages) and upstreamQuality (mean checker score of consumed inputs). Aggregation is like-for-like only — billing@d1 never pools with billing@d4 — and the short templates guarantee end-anchored types also occur shallow. The delta between a worker's isolated matrix score and its in-chain score at depth is published as the context tax, a first-class number rather than a hidden bias.

8. A/B propagation: the downstream effect of teaching

runABPair runs one ticket twice with identical rotation; only the pivot stage's worker differs (e.g. plan by to-prd@skill vs to-prd@bare). Both arms share a pairId; the published quantities are the downstream per-stage checker deltas, coherence delta, and chain delta — the project-level effect of a teaching source, which a per-stage score cannot see. Standard pivots: plan (skill vs bare) and build (skill vs bare).

9. Reliability rules

  • Preflight, twice: before any batch, every roster entry must resolve, mount its teaching source, and ground a real cwd (ax agency preflight, aborts loudly); before any ticket, its concrete assignments are re-verified — no chain spends a token on a worker that cannot start.
  • Never fake: unavailable judges are skipped; unmeasured is not zero; failed stages are recorded with their errors, and infra failures (ENOSPC) are retried rather than published as tool failures.
  • Hygiene: grounding templates are freed immediately after mount; aged trial dirs are swept during batches.

10. Versioning

This document is hash-locked (docs/methodology-agency.lock); every agency run manifest carries {version, hash} in provenance. A hash mismatch between runs means they were produced under different rules and are not comparable.

11. Known limitations (v1.0)

  • Stage corpus is 10 types across ~15 categories; video/audio/CAD/electronics stage types are designed but not yet wired.
  • Single universe → domain bias unquantified until a second client exists.
  • n per (worker × stage × depth) cell is small at current volume; chain-level findings harden faster than worker-level ones.
  • Judge panel is two models; disagreement is recorded, medians need volume.
  • Coherence covers what is grep-able; semantic contradictions between deliverables are only caught by the bundle judge.