OverGrid AI Field Notes · current study + historical archive

Good game AI begins with honest measurement.

A living record of deterministic self-play: what ran, what failed, what improved, and why no horizon snapshot is quietly promoted into a win.

62 analyzed studies133 arena runs38,621 match exposures1.04B candidate evaluations
2.11Msimulated turns
150,468player-seat exposures
30 maps172 seeds · 2–6 players
97.1 hsummed solver time
210.3 MiBcompressed raw evidence
37.3% territory complete8.6% flip-war stop54.1% horizon cappedMixed protocols: short diagnostic caps and full games are pooled here.

These are executed exposures across 41 research families, not 27,003 independent games: paired controls and condition reruns intentionally reuse 11,618 local match IDs. Zero recorded match failures does not include commands that stopped before writing an artifact.

Loading evidence…

Loading evidence…

Loading evidence…

Loading evidence…

Loading evidence…

Loading evidence…

Loading evidence…

Scroll to load historical Ender v5 trajectory

Scroll to load named-opponent calibration

Scroll to load historical Autoresearch v1

Research report · living edition

Learning OverGrid by teaching machines to play it

Abstract. We use deterministic self-play as both an engineering test and a strategy microscope. The goal is not merely to maximize a win counter: it is to build opponents that are legal, responsive, measurably strong, recognizably different, and useful for understanding the game. The program has produced reliable geometry findings, a practical search-scaling picture, safer endgame computation, and a much stricter definition of what counts as an AI improvement.

Archive trajectory

Autoresearch field log: improvement, corrections, and compute over time

The blue frontier uses one comparable KPI only: lower average score gap versus the frozen Phase 8 group panel. Later protocols answer different questions, so the line stops rather than pretending every experiment shares a scoreboard.

Swipe to read the full field log →BEST-SO-FAR GROUP SCORE-GAP IMPROVEMENTPHASE 8 · BUDGET 400 · CAPPED-HORIZON DIAGNOSTIC, NOT WIN RATE0+1+2+3+4+5Iteration 0: Established all-profile height-route diagnostics on every official map; actual improvement 0.00, best so far 0.00Iteration 1: Increased Raider height pressure while preserving aggression; actual improvement 1.99, best so far 1.99Iteration 4: Mountaineer completion transfer: height -0.06, target progress +0.06; natural finishes up, rank and summit identity down; actual improvement 1.01, best so far 1.99Iteration 5: Collective specialization: Balanced completion, Mountaineer completion, Builder fortification; actual improvement 4.25, best so far 4.25Iteration 6: Named randomness tax: Alai variety 20 to 12, Petra 7 to 5; both stronger, only Petra preserves style separation; actual improvement 2.53, best so far 4.25Iteration 7: Combined Pareto roster: Balanced and Mountaineer completion, Builder fortification, Petra precision; three finishes but Ender and Raider suppressed; actual improvement -0.60, best so far 4.25Iteration 8: Narrow Pareto roster: Builder fortification plus Petra precision; both stronger, closed-roster peers lose relative rank; actual improvement -1.88, best so far 4.25baseline+1.99+4.25PROTOCOL RESEToutcomes and schedules changelast comparable KPI · not remeasuredRESEARCH PROCESS · EVERY FAILED BRANCH REMAINS IN THE LOGJul 19 · 8Frozen group baselineComparable KPI beginsJul 19 · 8Portfolio best−4.25 score-gap pointsJul 20 · 12Final-round correctionEngine outcome resetJul 20 · 20Seat bundle bias foundRaw wins rejected as KPIJul 20 · 24L9 screen invalidatedNamespace confoundJul 22 · 25Selective search retainedUnsafe finishes −36.5ppJul 22 · 27/29Two ending rules rejectedCoverage ↔ fairnessJul 23 · 31/32No auto-promotionClosure + doctrine diagnosticsJul 23 · DeepMechanisms replicated480-game position screenCUMULATIVE CANDIDATE CONSTRUCTIONS5.8M20.4M203M671M702M1.04BAccepted, rejected, and descriptive arena work all count. Blocked and offline screens add zero arena compute.retained evidencerejected branchdiagnostic onlyengine correction
Reading. The first improvement frontier is real but narrow: +4.25 score-gap points on one capped-horizon group protocol. The more important later gains are methodological and capability-specific—fairer schedules, valid causal pairing, safer finishes, and replicated strategy mechanisms. They are annotated, not silently converted into the early KPI.
Context

What kind of game is OverGrid?

Every move places one connected three-dimensional piece. Its footprint claims the cells beneath it. A claim becomes a durable capture when the surrounding positions are controlled; opponents may cover soft claims but must bridge over solid captures. Dotted height tiles require an exact stack height and award their bonus only when captured.

This makes a move do several jobs at once: expand reach, harden territory, pursue height, cross protected cells, steal soft ownership, deny an opponent, and preserve shapes for future hands. The central research problem is therefore positional conversion: which visible progress actually becomes banked score before another player interrupts it?

Research subjects

The named AIs are competing hypotheses about good play

All named opponents see the same public board and use the same responsive candidate budget. Their identity comes from how they break strength-safe ties and which coherent plan they keep for a whole game.

easy opponent

Alai

identity designed · signature unresolved

Can optionality become a real strategic advantage?

Doctrine
Preserve multiple credible lanes and avoid committing the whole position through one fragile hinge.
Development
Random-looking variety became seeded, whole-game plans. The current doctrine may break only exact strength ties.
Evidence so far
Mobility sometimes rises, but independent-route superiority and multiplayer strength remain unproved.
medium opponent

Petra

clearest measured specialty

Can a bot turn height routes into reliable banked points?

Doctrine
Prefer short, defensible routes that convert exact-height access into capture, not merely motion toward a peak.
Development
A simple height bias evolved into route regret, conversion clocks, and compact-ascent plans.
Evidence so far
The clearest specialty so far: early held-out screens improved conversion, and Phase 32 reduced height-route regret.
hard opponent

Ender

strong baseline · doctrine constrained

Can reply search identify and deny the decisive opponent plan?

Doctrine
Use complete public reply search when available; otherwise pressure the actual leader without abandoning progress.
Development
Generic leader damage became winner-aware terminal evaluation, reply search, and targeted-denial plans.
Evidence so far
A strong duel benchmark, but Phase 32 lost score margin without establishing a clearer denial signature.
Wild opponent

Bean

distinct theft style · volatility unresolved

Can calculated risk create genuine high-upside play?

Doctrine
Attack thin seams and accept exposure only when the available score, height, or theft swing is unusually large.
Development
Move-level chaos became coherent jackpot, infiltration, and bailout-aware whole-game plans.
Evidence so far
Theft is visible; broader outcome volatility is not. “Wild” remains a play style, not a proven strength claim.
Why four generic controls remainThey test whether a named story adds value beyond a simple priority vector.
BalancedCommon reference

Score, height, structure, pressure, and mobility without one dominant specialty.

MountaineerHeight control

A simpler height-forward specialist used to test whether named Petra adds more than weight.

BuilderStructure control

Compact territory, capture chains, and future launch geometry.

RaiderPressure control

Opponent damage and leader pressure, with less emphasis on positional stability.

Development history

How the research question changed

1Stages 0–7

Make the arena trustworthy

We built reducer-backed headless play, deterministic replays, legal move checks, two-ply reply search, multiplayer seating, and separate live-versus-headless timing. Early win tables were diagnostics, not conclusions.

2Stages 8–22

Stop optimizing a broken scoreboard

Group autoresearch exposed turn caps, seat and starter bias, misleading horizon leaders, flip wars, and the difference between changing play style and improving strength. Controls became paired and interaction-balanced.

3Stages 23–30

Separate strength, style, and finishing

Named doctrines received explicit metrics. Orthogonal screens varied the whole roster together. Selective endgame search improved safety, while several plausible termination and doctrine changes were rejected.

4Stages 31–32 + deep screen

Measure mechanisms, not mythology

Neutral closure and named doctrines passed measurement gates but earned no automatic promotion. A 480-game position census established several board-geometry mechanisms while explicitly withholding win-rate causality.

Primary findings

What we think we know—and what each result does not prove

01
replicated mechanism

The best-supported strategy is geometric efficiency with an escape route.

Claim

Touch existing territory lightly, stay flat unless height buys something now, and compact the position without sealing every future launch face.

Evidence

Across 480 games and 120 map×seed clusters, minimal contact added 0.68 neutral columns, flat no-payoff placements added 1.62, purposeful height added 6.46 immediate points, and compact moves added 1.19 near-closures.

Interpretation

Compactness also reduced mobility by 4.16. “Grow a lobe, not a ribbon” is incomplete unless the lobe preserves useful exits.

Quoted evidence · strategy mechanism screenEfficient expansion has four distinct payoffs
Minimal contact
+0.68neutral columns
Flat, no payoff
+1.62neutral columns
Purposeful height
+6.46immediate points
Compact move
+1.19near-closures
Compact move
−4.16future placements

Mean paired move effect across 480 games; compactness buys closure but spends mobility.

Inspect supporting figures ↓
02
descriptive scaling

Search value converges faster than move identity.

Claim

More candidates help the evaluator find near-best moves, but the chosen move can keep changing after practical value is nearly flat.

Evidence

Mean regret to the 1,600-candidate reference fell from 0.0191 at 300 candidates to 0.0079 at 800. On deeper sentinels, 1,600 and 3,200 chose the same best move only 54.6% of the time, yet residual regret averaged 0.0064.

Interpretation

A live bot need not reproduce the deepest move exactly. It needs a latency budget where additional search buys a meaningful continuation advantage.

Quoted evidence · candidate censusEvaluator regret falls before move identity stabilizes

Lower is better. Reference regret reaches zero by definition at the 1,600-candidate census.

Inspect supporting figures ↓
03
paired safety result

Selective compute improved finish safety—not proven strength.

Claim

Spend extra search at high-information late-game decisions instead of slowing every turn.

Evidence

The held-out combined policy reduced unsafe finishes by 36.5 percentage points. Reserve search activated on 23.3% of measured calls and found a reserve-only completion on 6.5% of activations.

Interpretation

Completion, score, and competitive-strength intervals remained unresolved. This justifies an event-triggered reserve, not a claim that the AI became stronger.

Quoted evidence · paired endgame validationReserve search is rare, targeted, and safety-positive
−36.5 ppunsafe finishes
23.3%calls activating reserve
6.5%activations finding a reserve-only completion

This is an unsafe-finish result. It is not yet a demonstrated win-rate lift.

Inspect supporting figures ↓
04
equal-compute diagnostic

Personality is measurable only when it survives a strength guardrail.

Claim

A named bot should make recognizably different near-optimal choices, not receive hidden information, extra compute, or permission to play worse.

Evidence

In Phase 32, Petra clearly reduced height regret. Ender lost opponent-adjusted score margin without a clear denial gain; Alai mobility and Bean volatility intervals crossed zero.

Interpretation

Live named plans now break exact shared-strength ties only. The result is less theatrical than a loose style budget, but it keeps difficulty honest.

Quoted evidence · Phase 32 equal-compute diagnosticOnly one score-margin interval excluded zero—and it was negative

Dots show opponent-adjusted score-margin change; bars are 95% intervals across three independent seed clusters.

Inspect supporting figures ↓
05
methodological finding

The game engine and experiment design are part of the scientific instrument.

Claim

Natural wins, censored leads, adjudicated flip wars, and replay failures must remain separate outcomes.

Evidence

38,621 archived match exposures include repeated controls; 54% ended at mixed diagnostic or game horizons. Later studies added balanced starters, interaction-balanced seats, equal-compute counters, and exact command replay.

Interpretation

Several apparent AI regressions were measurement changes or objective failures. A bigger arena run cannot repair a biased schedule or a value function that rewards ending while behind.

Quoted evidence · archive-wide termination ledgerMost archived exposures are diagnostics, not completed games
44.2% territory complete0.7% flip-war stop55.1% declared horizon

Mixed protocols are pooled here. The cap share describes the research archive, not production game quality.

Inspect supporting figures ↓
Reading the evidence

Terms that prevent a plausible chart from becoming a false claim

Match exposure
One executed game condition. Paired controls and treatment reruns count separately, so exposures are not independent games.
Natural completion
Ordinary territory triggered the final round and all owed replies were completed under the game rules.
Horizon cap
The study stopped at its declared observation limit. The leader is recorded, but is not relabeled as a winner.
Candidate construction
One legal piece placement considered by the deterministic solver. Different constructions can resolve to the same board outcome.
Shared-strength gate
The common evaluator ranks moves first. A personality doctrine may choose only within the permitted strength-loss band; live named bots currently use an exact-tie band.
Doctrine
A character-specific preference among strength-safe moves, such as Petra preferring reliable height conversion or Alai preserving optionality.
Paired result
Control and treatment use the same map, seed, seats, starter, hands, and compute contract so their difference is interpretable.
Score AUC
Area under the score-over-time curve. It distinguishes sustained advantage from a late snapshot, but it is not a win rate.
Working conclusion

The ideal OverGrid AI is not simply the one that searches longest

It should find strong moves under a responsive budget, reveal a coherent strategic doctrine, remain fair under identical public information, and become more useful to a human observer as it improves.

The strongest program-wide lesson is that search and evaluation cannot be separated. Search tells us how effectively the machine explores the move space; the evaluator tells it what “good” means. When the evaluator confuses ending with winning, activity with conversion, or compactness with flexibility, extra compute magnifies the mistake.

The best current architecture is therefore shared strength plus bounded character: one public-board observation layer, one legality and search contract, character doctrines that operate only among strength-safe moves, seeded plans that stay coherent for a whole game, and selective reserve compute at decisions where more information can change the outcome.

For players, the machines are already teaching something general: efficient reach matters, verticality needs a concrete payoff, compactness needs exits, and a route matters only when it converts. The next research phase should test whether those one-move mechanisms survive into complete-game advantage.

Proposed additions

What would make this a more complete research paper?

Formal benchmark suite

Freeze representative opening, midgame, endgame, duel, and multiplayer positions with declared promotion gates.

Failure atlas

Show annotated board states where each AI made an obviously bad move, what feature misled it, and whether search or evaluation was at fault.

Human comparison

Compare AI move preferences with expert and beginner choices, including which AI explanations actually teach useful strategy.

Generalization tests

Hold out maps, player counts, seeds, and rule variants so improvements cannot specialize to the development panel.

Ablation table

Remove one observation, doctrine tool, reserve trigger, or evaluator term at a time to identify what actually causes an improvement.

Release and reproducibility record

Connect every live AI version to its solver fingerprint, research decision, artifact set, latency envelope, and rollback point.

↑ Figure 1 · current Ender hillclimb

Figures and diagnostic tables

The prose report above states the conclusions and their boundaries. Open a figure set to inspect effect sizes, intervals, protocol details, and subgroup reversals.

Figures 2–6Strategy mechanisms and deep position screenGeometry · selective compute · completion safety · 480 source games
Figures 7–8Historical collective interventionsPre-v8 neutral closure · named doctrines · strength and style kept separate
Figures 9+Detailed trajectory and scaling diagnosticsIteration history · compute curves · multiplayer response · collective controls
Research archiveHistorical research log4 eras · stages 0–30 · July 19–23, 2026