July 28 correction: the original doctrine baseline raised this concern; the later 61-state identical-public-position audit made it decisive. This retrospective note keeps both facts together and dates the research epoch to the baseline.

The embarrassing version is also the useful one: we had given Alai, Petra, Ender, and Bean names, biographies, difficulty labels, and supposedly distinct strategic instincts. Then we put them in the same public positions and asked a rude but necessary question:

Do they actually choose different constructions?

On the first 61-state identical-public-position corpus, the answer was no. Every incumbent profile selected the same move in every tested state.

That is not a tiny regression. It is a category error. A character who only gets a different label after the shared evaluator has already selected the move is not a game opponent with a style. It is a very expensive tooltip.

The architecture that made the placebo

All four named profiles had an equal 800-candidate budget, the same legal-move enumeration, the same selective reply search, and the same shared strength evaluator. The “character” layer acted only inside an extremely narrow near-best band. In practice that made it an exact-tie breaker far more often than a decision-maker.

That rail was originally a reasonable safety device: no character should become a covert rules exception, read hidden future deals, or casually throw away a winning move to sound in-character. The mistake was treating the safety rail as a personality engine.

What we saidWhat the first public-state audit foundWhat that actually means
Alai preserves options0 / 61 different incumbent movesno measured optionality policy was selecting
Petra converts height0 / 61 different incumbent movesno distinct route tool was selecting
Ender controls replies0 / 61 different incumbent movesno opponent-control mechanism was selecting
Bean seeks upside0 / 61 different incumbent movesno tail policy was selecting

A much larger baseline gave the same warning, in a subtler form

The historical, pre-fairness named-doctrine baseline ran 1,584 scheduled games across two-, four-, and six-player settings, official Standard and Advanced maps, balanced seating and starters, and 12 independent gameplay seeds. It was not a toy sample. Its point was to separate strength from identity.

The historical panel suggested Ender led the named two-player slice at 0.700, Petra led the multiplayer slice, Alai fell as tables grew, and Bean stole territory aggressively. These were hypotheses, not retained strength rankings: later home/rotation and terminal audits invalidated the panel for that use.

But the advertised stories did not all survive their own measurements:

ProfileWhat genuinely showed upWhat did not yet earn the copy
Alaihigh mobility and option entropy in two-player playbroad seeded variety was doing much of the visible work, and mobility vanished at six players
Petra+12.8 / +13.8 / +11.3pp height-conversion advantages in 2P / 4P / 6Pnothing: this was the first named identity with a clear behavioral receipt
Enderprecise, low-randomness duelingdemonstrated comeback control or a real opponent model
Beanthe most hard theft per turnthe promised high-risk, high-reward payoff distribution

The weirdest result was Bean. A squared positive-swing weight sounded “wild,” but its immediate swing width was lower than peers in every table size. The policy was reliably finding big moves, not a meaningful upside/downside distribution. Bean was better described as a thief than a gambler.

Strength, identity, and difficulty are three different knobs

Annotated reasoningStrength, identity, and difficulty require different evidence.

Read left to right. Each card names the claim it is allowed to support.

  1. StrengthDoes it win?

    Compare resolved outcomes at equal legal supply and compute.

    Outcome evidence
  2. IdentityDoes it choose differently?

    Trace the first public-state decision changed by the named doctrine.

    Action separation
  3. DifficultyHow forgiving is it?

    Vary acceptable whole plans without injecting per-move noise.

    Player experience
  4. PlaceboDoes it only sound different?

    Different copy over the same action trajectory is characterization, not policy identity.

    Do not promote
Read this asVoice can describe a policy after it acts. Only a public mechanism that changes the selected move establishes behavioral identity.

The repair was conceptual before it was technical.

  1. Strength asks whether a policy improves game outcomes under equal compute.
  2. Identity asks whether it reliably selects a different near-equal plan for a public board reason.
  3. Difficulty asks how widely it samples among acceptable plans.

Those cannot share one unlabeled “personality weights” knob. A competent profile needs a shared legal, deterministic, public-information strength rail. A named profile then needs at least one genuine specialization: a derived observation, a candidate-generation lens, a response model, or a selection mechanism that can change the shortlist. Difficulty variety belongs only after that, at the whole-plan level rather than as per-move dice rolling.

The research question that replaced “make them more different”

The next question was more precise: when a named policy sees an opportunity, does it activate, choose a different action, survive the opponent’s reply, and convert the promised advantage?

That creates an honest funnel:

opportunity → activation → first different move → reply survival → conversion → held-out result

It also makes failure informative. An inert policy is not “almost working.” A policy that changes moves but loses on Bigger is not “personality.” It is a useful hypothesis with a boundary.

That is the standard future posts will use. The characters get to be vivid. The evidence has to be boringly specific.

For a player’s version of inspectable risk, read Bean’s fork: two threats, one exit.