Rote Playoffs · WeMakeDevs × Modiqo · measured 2026-09-05

Procedure memory
beats model effort.

Same task. Same agent, same model, same prompt. The only difference: one arm was handed a crystallized Play — a recorded working method — and the other had to re-derive it from scratch. Ten trials, frozen ground truth, a deterministic grader, raw telemetry, and one answer key that the baseline agents found and used. Everything is on this page.

0×
faster wall-clock
36.7 s vs 362.7 s mean
0×
fewer output tokens
1,232 vs 18,905 mean
~0×
lower billed cost
$0.016 vs $0.154 mean
0
Plays published · all 1.00
zero API keys · read-only
01 · The scoreboard

The measurement, per trial.

Harness: Grok Build headless (grok-4.6-build), n=3 per arm, back-to-back, fresh scratch per trial. Correctness 0–10 against frozen ground truth via a deterministic grader. Hostile-input control (n=1 per arm) uses fixture-local truth. Full raw telemetry is in the repo — the table below is the complete result set.

36.7 s vs 362.7 s
mean wall-clock, with-Play vs baseline (n=3). With the anti-shortcut guard, honest re-derivation took 800–1,058 s — 25× slower than the Play.
1,232 vs 18,905
mean output tokens per run. The with-Play arm resolves a URI, runs it, reads the answer. The baseline writes an analyzer.
9 / 9 / 9
with-Play correctness, three trials, zero variance. Under corrupted input it reproduced ground truth exactly; the baseline diverged (3/10).
trialarmscore /10wall sturnsinput tokcache tokoutput tokcost
1baseline10*49922146,4551,410,04825,711$0.196
2baseline10*2581993,173982,78412,979$0.128
3baseline933118161,981754,30418,026$0.138
2baseline_v2180033————
3baseline_v211,05820————
1with_play933412,70775,9041,163$0.012
2with_play945524,02194,4641,589$0.018
3with_play932434,95949,920945$0.017

* leaked trials — see below; perfect by construction, kept and reported. baseline_v2 trial 1 exceeded the 900 s cap with no answer. Cost as billed by the harness; arm means: baseline 9.7 pts / $0.154 · with-Play 9.0 pts / $0.016.

hostile input

We fed both arms a corrupted transcript store — valid sessions mixed with a corrupt-JSONL file and one 600KB line. The with-Play arm reproduced the fixture truth exactly (single inert file, 1 read, 3 sessions, skips labeled — 10/10). The baseline degraded gracefully (no fail-open) but diverged — it missed the valid signal inside the partially-corrupt file (3/10). Labeled degradation beats plausible guessing.

Replicated across six tasks.

Five transfer tasks, each exercising a different Play and a different ability — side-effect auditing, instruction-burial analysis, commit hygiene, JSON integrity, live event extraction. Twenty further trials. Hardened protocol: baseline inputs staged inside the scratch, explicit anti-leak rules, every trial audited (zero leaks). Graded 0–10 against per-task frozen ground truth.

task (ability tested)with-Play, n=2baseline, n=2what separates them
side-effect audit — claim-vs-observation join10, 107, 7baseline misses the play's exact path-mention semantics
instruction burial — rules past the attention cliff9, 9 @ ~32 s2, 2 @ 572–2,180 sone baseline run spent 36 minutes deriving a wrong answer
commit hygiene — convention classification10, 108, 8baseline invents a looser convention
JSON integrity — error localization10, 1010, 9method is well-known; re-derivation is cheap here
event deadline — live multi-source join10, 1010, 10baseline correct but 3–4× slower per run

with-Play mean across the transfer suite: 9.8/10 (min 9). The boundary is the finding: procedure memory adds nothing where the method is common knowledge — and is decisive everywhere else.

02 · The finding nobody planned

Your agent will find the answer key.

leaked

Two of three baseline trials scored 10/10 — because the agents explored the filesystem, found this experiment's own frozen ground truth and reference engine, and used them. Their logs say it out loud:

"I'll run the official analyzer and check it against the frozen ground truth."

We kept the trials, marked them leaked, and amended the protocol mid-execution. Then we ran a hardened arm told to derive everything and reuse nothing — it took 800–1,058 seconds per trial and still answered wrong, under-scanning the transcripts (8 of 13 sessions) and inventing a weaker influence criterion. On a shared filesystem, benchmark isolation isn't a nicety. It's the experiment.

03 · How we ran it

Design, frozen before execution.

One real job — audit an AI agent's own JSONL session transcripts for files loaded into context that never influenced later work. Two transcript schemas, path extraction, order-sensitive influence tracing: genuinely nontrivial to re-derive, deterministic to grade.

Arms

with-Play — one added line: the exact rote play run command for the job's Play. It still has to run it and read the output. baseline — identical prompt, no mention of rote. baseline_v2 — plus an anti-shortcut guard, added after the leak.

n=3 per arm · hostile control n=1

Grading

Frozen ground truth from a reference implementation; a fixed Python grader scores free-text answers: top-3 set match (0–5), read counts ±1 (0–2), session count ±1 (0–2), influence semantics (0–1). Path matching is normalized to favor the baseline.

score_answers.py · offline

Honesty hardware

Every trial's text + reasoning grepped for leak patterns; results annotated per trial. Protocol amended twice, in writing, before grading. Telemetry, prompts, ground truth, and the grader are all published with the study.

10 trials · 0 missing datapoints
04 · The improvement loop

Three real bugs, caught by real runs.

Rule for every update: score 1.00 → push → download the published Play fresh → run it on something real. No demos. The loop caught what staring at code never would:

buried-instructions 0.1.1

A detector that couldn't read real instructions

Run on the actual AGENTS.md: 0 rules found — in a file full of rules. Markdown emphasis broke the matcher and rules hid mid-paragraph. Rewritten sentence-level with an imperative-verb pattern: now 303 rules across 21 real skill files, 88 past the ~3,000-token reading cliff.

catch-sneaky-changes 0.2.1

A census that could lie on big repos

The after-census flagged files as "dirty→cleared" that still existed — an artifact of the 500-entry git cap shifting alphabetically between snapshots. Fixed: cleared-path claims suppressed under truncation, warning surfaced. Re-verified: 16/16 real transitions, zero artifacts. Now a 4-probe parallel DAG (lockfiles ∥ listeners ∥ git + join).

find-wasted-tokens 0.2.0

Scope before you scan

Upgraded to a parallel pair: a no-parse inventory census (∥) the deep influence analysis, so you see what will be scanned — 13 sessions, 12.4 MB on this machine — before committing to the run. Ran live on real transcripts; the published URI is what you can run below.

05 · The arsenal

38 Plays. Every one scored 1.00.

Zero API keys. Python stdlib only. Read-only unless the contract discloses writes. Inspectable before you run it — rote play inspect the URI first. Standing after the build week: #4 of 123 playmakers in a catalog of 800+ Plays.

06 · The paper

"Procedure Memory Beats Model Effort."

A controlled case study of one real task — auditing an AI agent's own session transcripts for context files that were loaded but never influenced later work — executed by the same harness in two arms: with a crystallized Play, and re-deriving from scratch. The with-Play arm finished in 36.7 s on average versus 362.7 s (9.9×), used 15.3× fewer output tokens at ~10× lower billed cost, and exactly reproduced ground truth under hostile input where the baseline diverged — at equal correctness on clean inputs. Two of three baseline trials were leaked: they discovered the experiment's own ground truth on the shared filesystem and used it, scoring perfectly by construction; we keep, report, and learn from them. The honest baseline's failure modes — the method-rebuilding tax, partial-corruption divergence, and the price of anti-shortcut guarding — are catalogued with session receipts. Full protocol, prompts, telemetry, and grader ship with the study.

Contribution 1

A paired-arm protocol (with-Play vs baseline) for measuring procedure reuse: frozen ground truth, deterministic grader, leak audit, pre-registered amendments.

Contribution 2

Per-trial results on a real task with a hostile-input control — including a benchmark-isolation failure every agent-evaluation builder should internalize.

Contribution 3

A baseline failure-mode taxonomy: answer-key capture, the method-rebuilding tax, fails-plausible degradation, and the cost of forced honesty.

Manual arms (documented, runnable prompt packs included): Gemini 3.8 Flash High in Google Antigravity 2.0 — $0.75/$3.75 per 1M tokens, adjustable thinking depth, strong agentic-coding gains; Meta Muse Spark 1.3 in OpenCode Zen — free, #4 by observed usage (~10T tokens/week, 93% input-cache ratio). The thesis predicts the largest savings exactly where models are cheapest.

07 · Run it yourself

One command. Then one URI.

Credentials never travel; the contract declares every effect. This is the whole submission — no deck, no video, just a public URI that runs.

curl -fsSL https://getrote.dev/playoffs/install.sh | sh
rote play run https://play.modiqo.ai/bhanu82/find-wasted-tokens