RESEARCH NOTES · NAMED OPPONENTS
Behind the AIFour named AIs walked into the same position. They all made the same move.
July 28 correction: the original doctrine baseline raised this concern; the later 61-state identical-public-position audit made it decisive. This retrospective note keeps both facts together and dates the research epoch to the baseline.
The embarrassing version is also the useful one: we had given Alai, Petra, Ender, and Bean names, biographies, difficulty labels, and supposedly distinct strategic instincts. Then we put them in the same public positions and asked a rude but necessary question:
Do they actually choose different constructions?
On the first 61-state identical-public-position corpus, the answer was no. Every incumbent profile selected the same move in every tested state.
That is not a tiny regression. It is a category error. A character who only gets a different label after the shared evaluator has already selected the move is not a game opponent with a style. It is a very expensive tooltip.
The architecture that made the placebo
All four named profiles had an equal 800-candidate budget, the same legal-move enumeration, the same selective reply search, and the same shared strength evaluator. The “character” layer acted only inside an extremely narrow near-best band. In practice that made it an exact-tie breaker far more often than a decision-maker.
That rail was originally a reasonable safety device: no character should become a covert rules exception, read hidden future deals, or casually throw away a winning move to sound in-character. The mistake was treating the safety rail as a personality engine.
| What we said | What the first public-state audit found | What that actually means |
|---|---|---|
| Alai preserves options | 0 / 61 different incumbent moves | no measured optionality policy was selecting |
| Petra converts height | 0 / 61 different incumbent moves | no distinct route tool was selecting |
| Ender controls replies | 0 / 61 different incumbent moves | no opponent-control mechanism was selecting |
| Bean seeks upside | 0 / 61 different incumbent moves | no tail policy was selecting |
A much larger baseline gave the same warning, in a subtler form
The historical, pre-fairness named-doctrine baseline ran 1,584 scheduled games across two-, four-, and six-player settings, official Standard and Advanced maps, balanced seating and starters, and 12 independent gameplay seeds. It was not a toy sample. Its point was to separate strength from identity.
The historical panel suggested Ender led the named two-player slice at 0.700, Petra led the multiplayer slice, Alai fell as tables grew, and Bean stole territory aggressively. These were hypotheses, not retained strength rankings: later home/rotation and terminal audits invalidated the panel for that use.
But the advertised stories did not all survive their own measurements:
| Profile | What genuinely showed up | What did not yet earn the copy |
|---|---|---|
| Alai | high mobility and option entropy in two-player play | broad seeded variety was doing much of the visible work, and mobility vanished at six players |
| Petra | +12.8 / +13.8 / +11.3pp height-conversion advantages in 2P / 4P / 6P | nothing: this was the first named identity with a clear behavioral receipt |
| Ender | precise, low-randomness dueling | demonstrated comeback control or a real opponent model |
| Bean | the most hard theft per turn | the promised high-risk, high-reward payoff distribution |
The weirdest result was Bean. A squared positive-swing weight sounded “wild,” but its immediate swing width was lower than peers in every table size. The policy was reliably finding big moves, not a meaningful upside/downside distribution. Bean was better described as a thief than a gambler.
Strength, identity, and difficulty are three different knobs
Read left to right. Each card names the claim it is allowed to support.
- StrengthDoes it win?
Compare resolved outcomes at equal legal supply and compute.
Outcome evidence→ - IdentityDoes it choose differently?
Trace the first public-state decision changed by the named doctrine.
Action separation→ - DifficultyHow forgiving is it?
Vary acceptable whole plans without injecting per-move noise.
Player experience→ - PlaceboDoes it only sound different?
Different copy over the same action trajectory is characterization, not policy identity.
Do not promote
The repair was conceptual before it was technical.
- Strength asks whether a policy improves game outcomes under equal compute.
- Identity asks whether it reliably selects a different near-equal plan for a public board reason.
- Difficulty asks how widely it samples among acceptable plans.
Those cannot share one unlabeled “personality weights” knob. A competent profile needs a shared legal, deterministic, public-information strength rail. A named profile then needs at least one genuine specialization: a derived observation, a candidate-generation lens, a response model, or a selection mechanism that can change the shortlist. Difficulty variety belongs only after that, at the whole-plan level rather than as per-move dice rolling.
The research question that replaced “make them more different”
The next question was more precise: when a named policy sees an opportunity, does it activate, choose a different action, survive the opponent’s reply, and convert the promised advantage?
That creates an honest funnel:
opportunity → activation → first different move → reply survival → conversion → held-out result
It also makes failure informative. An inert policy is not “almost working.” A policy that changes moves but loses on Bigger is not “personality.” It is a useful hypothesis with a boundary.
That is the standard future posts will use. The characters get to be vivid. The evidence has to be boringly specific.
For a player’s version of inspectable risk, read Bean’s fork: two threats, one exit.