Learning OverGrid by teaching machines to play it
Abstract. We use deterministic self-play as both an engineering test and a strategy microscope. The goal is not merely to maximize a win counter: it is to build opponents that are legal, responsive, measurably strong, recognizably different, and useful for understanding the game. The program has produced reliable geometry findings, a practical search-scaling picture, safer endgame computation, and a much stricter definition of what counts as an AI improvement.
Autoresearch field log: improvement, corrections, and compute over time
The blue frontier uses one comparable KPI only: lower average score gap versus the frozen Phase 8 group panel. Later protocols answer different questions, so the line stops rather than pretending every experiment shares a scoreboard.
What kind of game is OverGrid?
Every move places one connected three-dimensional piece. Its footprint claims the cells beneath it. A claim becomes a durable capture when the surrounding positions are controlled; opponents may cover soft claims but must bridge over solid captures. Dotted height tiles require an exact stack height and award their bonus only when captured.
This makes a move do several jobs at once: expand reach, harden territory, pursue height, cross protected cells, steal soft ownership, deny an opponent, and preserve shapes for future hands. The central research problem is therefore positional conversion: which visible progress actually becomes banked score before another player interrupts it?
The named AIs are competing hypotheses about good play
All named opponents see the same public board and use the same responsive candidate budget. Their identity comes from how they break strength-safe ties and which coherent plan they keep for a whole game.
Alai
identity designed · signature unresolvedCan optionality become a real strategic advantage?
- Doctrine
- Preserve multiple credible lanes and avoid committing the whole position through one fragile hinge.
- Development
- Random-looking variety became seeded, whole-game plans. The current doctrine may break only exact strength ties.
- Evidence so far
- Mobility sometimes rises, but independent-route superiority and multiplayer strength remain unproved.
Petra
clearest measured specialtyCan a bot turn height routes into reliable banked points?
- Doctrine
- Prefer short, defensible routes that convert exact-height access into capture, not merely motion toward a peak.
- Development
- A simple height bias evolved into route regret, conversion clocks, and compact-ascent plans.
- Evidence so far
- The clearest specialty so far: early held-out screens improved conversion, and Phase 32 reduced height-route regret.
Ender
strong baseline · doctrine constrainedCan reply search identify and deny the decisive opponent plan?
- Doctrine
- Use complete public reply search when available; otherwise pressure the actual leader without abandoning progress.
- Development
- Generic leader damage became winner-aware terminal evaluation, reply search, and targeted-denial plans.
- Evidence so far
- A strong duel benchmark, but Phase 32 lost score margin without establishing a clearer denial signature.
Bean
distinct theft style · volatility unresolvedCan calculated risk create genuine high-upside play?
- Doctrine
- Attack thin seams and accept exposure only when the available score, height, or theft swing is unusually large.
- Development
- Move-level chaos became coherent jackpot, infiltration, and bailout-aware whole-game plans.
- Evidence so far
- Theft is visible; broader outcome volatility is not. “Wild” remains a play style, not a proven strength claim.
Score, height, structure, pressure, and mobility without one dominant specialty.
A simpler height-forward specialist used to test whether named Petra adds more than weight.
Compact territory, capture chains, and future launch geometry.
Opponent damage and leader pressure, with less emphasis on positional stability.
How the research question changed
Make the arena trustworthy
We built reducer-backed headless play, deterministic replays, legal move checks, two-ply reply search, multiplayer seating, and separate live-versus-headless timing. Early win tables were diagnostics, not conclusions.
Stop optimizing a broken scoreboard
Group autoresearch exposed turn caps, seat and starter bias, misleading horizon leaders, flip wars, and the difference between changing play style and improving strength. Controls became paired and interaction-balanced.
Separate strength, style, and finishing
Named doctrines received explicit metrics. Orthogonal screens varied the whole roster together. Selective endgame search improved safety, while several plausible termination and doctrine changes were rejected.
Measure mechanisms, not mythology
Neutral closure and named doctrines passed measurement gates but earned no automatic promotion. A 480-game position census established several board-geometry mechanisms while explicitly withholding win-rate causality.
What we think we know—and what each result does not prove
The best-supported strategy is geometric efficiency with an escape route.
Touch existing territory lightly, stay flat unless height buys something now, and compact the position without sealing every future launch face.
Across 480 games and 120 map×seed clusters, minimal contact added 0.68 neutral columns, flat no-payoff placements added 1.62, purposeful height added 6.46 immediate points, and compact moves added 1.19 near-closures.
Compactness also reduced mobility by 4.16. “Grow a lobe, not a ribbon” is incomplete unless the lobe preserves useful exits.
Mean paired move effect across 480 games; compactness buys closure but spends mobility.
Search value converges faster than move identity.
More candidates help the evaluator find near-best moves, but the chosen move can keep changing after practical value is nearly flat.
Mean regret to the 1,600-candidate reference fell from 0.0191 at 300 candidates to 0.0079 at 800. On deeper sentinels, 1,600 and 3,200 chose the same best move only 54.6% of the time, yet residual regret averaged 0.0064.
A live bot need not reproduce the deepest move exactly. It needs a latency budget where additional search buys a meaningful continuation advantage.
at 1,600 vs 3,200but only .0064 residual regret
Lower is better. Reference regret reaches zero by definition at the 1,600-candidate census.
Selective compute improved finish safety—not proven strength.
Spend extra search at high-information late-game decisions instead of slowing every turn.
The held-out combined policy reduced unsafe finishes by 36.5 percentage points. Reserve search activated on 23.3% of measured calls and found a reserve-only completion on 6.5% of activations.
Completion, score, and competitive-strength intervals remained unresolved. This justifies an event-triggered reserve, not a claim that the AI became stronger.
This is an unsafe-finish result. It is not yet a demonstrated win-rate lift.
Personality is measurable only when it survives a strength guardrail.
A named bot should make recognizably different near-optimal choices, not receive hidden information, extra compute, or permission to play worse.
In Phase 32, Petra clearly reduced height regret. Ender lost opponent-adjusted score margin without a clear denial gain; Alai mobility and Bean volatility intervals crossed zero.
Live named plans now break exact shared-strength ties only. The result is less theatrical than a loose style budget, but it keeps difficulty honest.
Dots show opponent-adjusted score-margin change; bars are 95% intervals across three independent seed clusters.
The game engine and experiment design are part of the scientific instrument.
Natural wins, censored leads, adjudicated flip wars, and replay failures must remain separate outcomes.
38,621 archived match exposures include repeated controls; 54% ended at mixed diagnostic or game horizons. Later studies added balanced starters, interaction-balanced seats, equal-compute counters, and exact command replay.
Several apparent AI regressions were measurement changes or objective failures. A bigger arena run cannot repair a biased schedule or a value function that rewards ending while behind.
Mixed protocols are pooled here. The cap share describes the research archive, not production game quality.
Terms that prevent a plausible chart from becoming a false claim
- Match exposure
- One executed game condition. Paired controls and treatment reruns count separately, so exposures are not independent games.
- Natural completion
- Ordinary territory triggered the final round and all owed replies were completed under the game rules.
- Horizon cap
- The study stopped at its declared observation limit. The leader is recorded, but is not relabeled as a winner.
- Candidate construction
- One legal piece placement considered by the deterministic solver. Different constructions can resolve to the same board outcome.
- Shared-strength gate
- The common evaluator ranks moves first. A personality doctrine may choose only within the permitted strength-loss band; live named bots currently use an exact-tie band.
- Doctrine
- A character-specific preference among strength-safe moves, such as Petra preferring reliable height conversion or Alai preserving optionality.
- Paired result
- Control and treatment use the same map, seed, seats, starter, hands, and compute contract so their difference is interpretable.
- Score AUC
- Area under the score-over-time curve. It distinguishes sustained advantage from a late snapshot, but it is not a win rate.
The ideal OverGrid AI is not simply the one that searches longest
It should find strong moves under a responsive budget, reveal a coherent strategic doctrine, remain fair under identical public information, and become more useful to a human observer as it improves.
The strongest program-wide lesson is that search and evaluation cannot be separated. Search tells us how effectively the machine explores the move space; the evaluator tells it what “good” means. When the evaluator confuses ending with winning, activity with conversion, or compactness with flexibility, extra compute magnifies the mistake.
The best current architecture is therefore shared strength plus bounded character: one public-board observation layer, one legality and search contract, character doctrines that operate only among strength-safe moves, seeded plans that stay coherent for a whole game, and selective reserve compute at decisions where more information can change the outcome.
For players, the machines are already teaching something general: efficient reach matters, verticality needs a concrete payoff, compactness needs exits, and a route matters only when it converts. The next research phase should test whether those one-move mechanisms survive into complete-game advantage.
What would make this a more complete research paper?
Freeze representative opening, midgame, endgame, duel, and multiplayer positions with declared promotion gates.
Show annotated board states where each AI made an obviously bad move, what feature misled it, and whether search or evaluation was at fault.
Compare AI move preferences with expert and beginner choices, including which AI explanations actually teach useful strategy.
Hold out maps, player counts, seeds, and rule variants so improvements cannot specialize to the development panel.
Remove one observation, doctrine tool, reserve trigger, or evaluator term at a time to identify what actually causes an improvement.
Connect every live AI version to its solver fingerprint, research decision, artifact set, latency envelope, and rollback point.