Same task. Same agent, same model, same prompt. The only difference: one arm was handed a crystallized Play — a recorded working method — and the other had to re-derive it from scratch. Ten trials, frozen ground truth, a deterministic grader, raw telemetry, and one answer key that the baseline agents found and used. Everything is on this page.
Harness: Grok Build headless (grok-4.6-build), n=3 per arm, back-to-back, fresh scratch per trial. Correctness 0–10 against frozen ground truth via a deterministic grader. Hostile-input control (n=1 per arm) uses fixture-local truth. Full raw telemetry is in the repo — the table below is the complete result set.
| trial | arm | score /10 | wall s | turns | input tok | cache tok | output tok | cost |
|---|---|---|---|---|---|---|---|---|
| 1 | baseline | 10* | 499 | 22 | 146,455 | 1,410,048 | 25,711 | $0.196 |
| 2 | baseline | 10* | 258 | 19 | 93,173 | 982,784 | 12,979 | $0.128 |
| 3 | baseline | 9 | 331 | 18 | 161,981 | 754,304 | 18,026 | $0.138 |
| 2 | baseline_v2 | 1 | 800 | 33 | — | — | — | — |
| 3 | baseline_v2 | 1 | 1,058 | 20 | — | — | — | — |
| 1 | with_play | 9 | 33 | 4 | 12,707 | 75,904 | 1,163 | $0.012 |
| 2 | with_play | 9 | 45 | 5 | 24,021 | 94,464 | 1,589 | $0.018 |
| 3 | with_play | 9 | 32 | 4 | 34,959 | 49,920 | 945 | $0.017 |
* leaked trials — see below; perfect by construction, kept and reported. baseline_v2 trial 1 exceeded the 900 s cap with no answer. Cost as billed by the harness; arm means: baseline 9.7 pts / $0.154 · with-Play 9.0 pts / $0.016.
We fed both arms a corrupted transcript store — valid sessions mixed with a corrupt-JSONL file and one 600KB line. The with-Play arm reproduced the fixture truth exactly (single inert file, 1 read, 3 sessions, skips labeled — 10/10). The baseline degraded gracefully (no fail-open) but diverged — it missed the valid signal inside the partially-corrupt file (3/10). Labeled degradation beats plausible guessing.
Five transfer tasks, each exercising a different Play and a different ability — side-effect auditing, instruction-burial analysis, commit hygiene, JSON integrity, live event extraction. Twenty further trials. Hardened protocol: baseline inputs staged inside the scratch, explicit anti-leak rules, every trial audited (zero leaks). Graded 0–10 against per-task frozen ground truth.
| task (ability tested) | with-Play, n=2 | baseline, n=2 | what separates them |
|---|---|---|---|
| side-effect audit — claim-vs-observation join | 10, 10 | 7, 7 | baseline misses the play's exact path-mention semantics |
| instruction burial — rules past the attention cliff | 9, 9 @ ~32 s | 2, 2 @ 572–2,180 s | one baseline run spent 36 minutes deriving a wrong answer |
| commit hygiene — convention classification | 10, 10 | 8, 8 | baseline invents a looser convention |
| JSON integrity — error localization | 10, 10 | 10, 9 | method is well-known; re-derivation is cheap here |
| event deadline — live multi-source join | 10, 10 | 10, 10 | baseline correct but 3–4× slower per run |
with-Play mean across the transfer suite: 9.8/10 (min 9). The boundary is the finding: procedure memory adds nothing where the method is common knowledge — and is decisive everywhere else.
Two of three baseline trials scored 10/10 — because the agents explored the filesystem, found this experiment's own frozen ground truth and reference engine, and used them. Their logs say it out loud:
"I'll run the official analyzer and check it against the frozen ground truth."We kept the trials, marked them leaked, and amended the protocol mid-execution. Then we ran a hardened arm told to derive everything and reuse nothing — it took 800–1,058 seconds per trial and still answered wrong, under-scanning the transcripts (8 of 13 sessions) and inventing a weaker influence criterion. On a shared filesystem, benchmark isolation isn't a nicety. It's the experiment.
One real job — audit an AI agent's own JSONL session transcripts for files loaded into context that never influenced later work. Two transcript schemas, path extraction, order-sensitive influence tracing: genuinely nontrivial to re-derive, deterministic to grade.
with-Play — one added line: the exact rote play run command for the job's Play. It still has to run it and read the output. baseline — identical prompt, no mention of rote. baseline_v2 — plus an anti-shortcut guard, added after the leak.
n=3 per arm · hostile control n=1Frozen ground truth from a reference implementation; a fixed Python grader scores free-text answers: top-3 set match (0–5), read counts ±1 (0–2), session count ±1 (0–2), influence semantics (0–1). Path matching is normalized to favor the baseline.
score_answers.py · offlineEvery trial's text + reasoning grepped for leak patterns; results annotated per trial. Protocol amended twice, in writing, before grading. Telemetry, prompts, ground truth, and the grader are all published with the study.
10 trials · 0 missing datapointsRule for every update: score 1.00 → push → download the published Play fresh → run it on something real. No demos. The loop caught what staring at code never would:
Run on the actual AGENTS.md: 0 rules found — in a file full of rules. Markdown emphasis broke the matcher and rules hid mid-paragraph. Rewritten sentence-level with an imperative-verb pattern: now 303 rules across 21 real skill files, 88 past the ~3,000-token reading cliff.
The after-census flagged files as "dirty→cleared" that still existed — an artifact of the 500-entry git cap shifting alphabetically between snapshots. Fixed: cleared-path claims suppressed under truncation, warning surfaced. Re-verified: 16/16 real transitions, zero artifacts. Now a 4-probe parallel DAG (lockfiles ∥ listeners ∥ git + join).
Upgraded to a parallel pair: a no-parse inventory census (∥) the deep influence analysis, so you see what will be scanned — 13 sessions, 12.4 MB on this machine — before committing to the run. Ran live on real transcripts; the published URI is what you can run below.
Zero API keys. Python stdlib only. Read-only unless the contract discloses writes. Inspectable before you run it — rote play inspect the URI first. Standing after the build week: #4 of 123 playmakers in a catalog of 800+ Plays.
A controlled case study of one real task — auditing an AI agent's own session transcripts for context files that were loaded but never influenced later work — executed by the same harness in two arms: with a crystallized Play, and re-deriving from scratch. The with-Play arm finished in 36.7 s on average versus 362.7 s (9.9×), used 15.3× fewer output tokens at ~10× lower billed cost, and exactly reproduced ground truth under hostile input where the baseline diverged — at equal correctness on clean inputs. Two of three baseline trials were leaked: they discovered the experiment's own ground truth on the shared filesystem and used it, scoring perfectly by construction; we keep, report, and learn from them. The honest baseline's failure modes — the method-rebuilding tax, partial-corruption divergence, and the price of anti-shortcut guarding — are catalogued with session receipts. Full protocol, prompts, telemetry, and grader ship with the study.
A paired-arm protocol (with-Play vs baseline) for measuring procedure reuse: frozen ground truth, deterministic grader, leak audit, pre-registered amendments.
Per-trial results on a real task with a hostile-input control — including a benchmark-isolation failure every agent-evaluation builder should internalize.
A baseline failure-mode taxonomy: answer-key capture, the method-rebuilding tax, fails-plausible degradation, and the cost of forced honesty.
Manual arms (documented, runnable prompt packs included): Gemini 3.8 Flash High in Google Antigravity 2.0 — $0.75/$3.75 per 1M tokens, adjustable thinking depth, strong agentic-coding gains; Meta Muse Spark 1.3 in OpenCode Zen — free, #4 by observed usage (~10T tokens/week, 93% input-cache ratio). The thesis predicts the largest savings exactly where models are cheapest.
Credentials never travel; the contract declares every effect. This is the whole submission — no deck, no video, just a public URI that runs.
curl -fsSL https://getrote.dev/playoffs/install.sh | sh
rote play run https://play.modiqo.ai/bhanu82/find-wasted-tokens