RESEARCH NOTES · AUTORESEARCH V1
Behind the AIOur first hillclimb had no finish line.
The first OverGrid autoresearch chart was exactly the kind of thing that makes an ML person sit up straighter: numbered experiments, a running best, discarded dots, compute, and tiny knobs that seemed to move a KPI.
It was also missing a finish line.
Read left to right. Each card names the claim it is allowed to support.
- InstrumentRun bounded games
Compare one small policy change at a fixed 40-turn horizon.
348 measured games→ - ObservationThe chart moves
Score, height, regret, and compute produce a tidy running best.
Useful diagnostic→ - Stop signEvery game is capped
No result reaches the rules-defined finish before measurement ends.
0 natural winners→ - VerdictKeep hypotheses, not rank
Archive the curve as research chronology and redesign the outcome contract.
No champion promoted
The graph was honest. The inference was not ready.
The original frozen screen ran 348 matches across 30 official maps, 2–6 players, eight profiles, balanced forward/reverse seats, a fixed seed, and 400 candidates per move. Every game was capped at 40 turns.
That makes it a real and useful diagnostic: it can show whether a change shifts height-route regret, score trajectory, rank proxies, style separation, or runtime. It cannot tell us that a policy wins complete games. A leader at a forced stop is a photograph of a race in progress, not a medal ceremony.
This is why the post carries an INVALID badge for strength inference, rather than pretending that a large n turns a censored outcome into a natural win rate. The games ran. The computation counts. The promotion claim does not.
What the hillclimb actually found
The loop was simple by design: one bounded change at a time; retain measured improvements; discard regressions; preserve the full ledger. Several outcomes were genuinely informative.
| Round | Hypothesis | What happened | Honest disposition |
|---|---|---|---|
| 2 | Alai should value height more | Regret fell, but p95 latency exceeded the guardrail | Discarded |
| 3 | Raider should press height more | Screen metrics improved | Kept provisionally |
| 4 | Bean should trade volatility for height | Screen rank improved; held-out natural wins fell 1→0 | Discarded |
| 5 | Ender should target the leader harder | Early behavior improved; longer holdout regressed | Discarded |
| 6–9 | Give different profiles narrower changes | Some Pareto branches emerged; other profiles lost relative rank | Research branches, not roster promotion |
The point of showing the discarded rows is not self-flagellation. It is to make the direction of travel legible. Bean’s row, for example, clarified that a generic “height regret” objective could make a high-variance profile less distinctive and remove the panel’s only natural win. That is a real lesson even though the candidate never shipped.
The first scaling curve: breadth helped, certainty did not
The next panel varied candidate budget and requested search depth. At 100, 400, and 1,600 candidates, higher root breadth improved control-relative score margin monotonically in the panel. The 1-ply group moved from −464.5 score margin at 100 candidates to −64.7 at 400 and +110.3 at 1,600.
That looks like a scaling law. It is better described as a scaling signal: the 192-game panel had only 13 natural finishes, four games per profile/depth/budget cell, and a 40-turn horizon. It tells us where to look for a compute knee. It does not tell us a globally optimal live budget or a definitive winner.
The more actionable result was mundane: candidate budget appeared to matter more reliably than requested depth. At the same root budget, a two-ply request sometimes helped, but it only realized depth two on about 54–61% of moves. “Depth two” on the chart was an intention, not a constant amount of calculation.
Hillclimbing is an instrument, not an oracle
Autoresearch did its intended job. It made every tweak explicit, preserved rejected work, counted compute, and made reversals visible. The later 10-round heads-up campaign made the same point even more sharply: a connected-height-route candidate showed a +0.84 discovery lead on three untouched seed clusters, then reversed to −0.35 in a 12-seed confirmation. No treatment was promoted.
That is not a failure of the method; it is the method refusing to flatter us.
<details> <summary>What changed after this epoch</summary>Later work separated natural outcomes from horizons, crossed starters and physical homes, repaired rotation fairness, isolated termination classes, and archived pre-fairness datasets for strength/ranking claims. The original curve remains valuable as a research chronology and a source of hypotheses, but it is not a valid leaderboard. The current audit trail and its newer evidence boundaries are in the AI Lab.
</details>The output we wanted was never “the chart went up.” It was a system that can tell the difference between a better policy, an easier opponent, a shorter game, a seat artifact, a stalled ending, and a seductive little line that happens to go northeast.
The discarded progress heuristic became a separate strategy note: A route is not a bank.
Test your understanding
No score, no account. Pick an answer and reveal the reasoning.
01All 348 games end at a fixed 40-turn horizon. What may the chart support?
02Why keep invalid experiments visible?