The dramatic early result did not survive the final test.

This page keeps the exploratory ranking, the separate finalist test, and the tactical audit visible. The distinction matters: the 100,000-game screen found possibilities; the untouched 24,000-game confirmation decided which claims could be published.

Total games
124,000
Exploration
100,000
Confirmation
24,000
Fresh deal clusters
1,000

No preregistered finalist comparison showed a meaningful advantage.

The four finalists ranged from 24.6% to 25.7% in descriptive win rate. Every preregistered difference included zero in its 95% clustered interval, and none reached the study’s three-percentage-point threshold.

Table Denial’s 8.11-point exploratory lead over Balanced became a 0.34-point deficit in confirmation. The likely lesson is context: a reactive strategy can thrive in one opponent ecology and look ordinary in another.

Three kinds of evidence

Tactical audit
Checks whether the engine and serious policies miss an immediate way to empty the rack. They missed none in 10,000 test positions.
Exploration
Tries many strategies and table mixes to find leads. Its rankings are useful for choosing the next test, not for declaring a winner.
Confirmation
Tests frozen questions on fresh deals. Only this stage can promote the preregistered comparisons to findings.
Percentage points
A direct win-rate change. Moving from 25% to 28% is a gain of three percentage points.

Not sure what a strategy name means? Read the plain-English strategy field guide first.

The four finalists finished close together

These are descriptive rates from the same cyclically balanced finalist table. Overlapping intervals mean the visible order should not be read as a proven ranking.

StrategyWin rate95% intervalMean scoreLosing rack
Greedy Tiles 25.7% 25.1% to 26.2% +0.42 28.30
Balanced 25.0% 24.5% to 25.6% +0.37 27.68
Table Denial 24.7% 24.2% to 25.2% -0.23 27.98
Opponent Modeler 24.6% 24.0% to 25.1% -0.56 28.17

What the final test actually asked

Each interval is clustered by the 1,000 independent deal blocks. A claim needed both an adjusted statistical result and an advantage of at least three percentage points.

ComparisonDifference95% intervalResult
Table Denial vs Balanced -0.34 points -1.06 to +0.39 Not supported
Table Denial vs Opponent Modeler +0.12 points -0.35 to +0.58 Not supported
Opponent Modeler vs Balanced -0.45 points -1.17 to +0.27 Not supported

Why Table Denial reached the final

The broad screen mixed homogeneous, two-versus-two, three-plus-one, and all-distinct tables. These results generated hypotheses. They do not override the held-out result above.

StrategyWin rateMean scoreRack value left
Table Denial 33.5% +13.59 17.21
Opponent Modeler 28.3% +6.11 19.17
Match Aware 26.9% +5.83 19.51
Strategic Drawer 26.7% -2.33 26.42
Manipulation Averse 26.6% -8.75 30.14
Greedy Tiles 26.1% +4.51 20.54
Shallow Planner 25.8% +4.82 18.97
Endgame Liquidator 25.8% +2.13 20.47
Greedy Points 25.8% +5.09 19.01
Eager Beginner 25.7% +4.01 20.50
Conservative 25.4% +3.78 19.89
Balanced 25.4% +3.60 19.45
Late Opening Hoarder 25.1% -7.89 30.67
Joker Holder 25.0% +3.64 18.93
Pivot Copy Keeper 24.9% -1.31 23.57
Joker Farmer 24.7% +0.28 20.60
Random Legal 4.4% -35.40 39.50

No single preference moved wins by even one point

A separate 32-combination experiment switched five tendencies on and off. All observed differences were smaller than 0.21 percentage points, another warning against one-rule strategy slogans.

Tendency emphasizedWin-rate differenceHigh settingLow setting
Hold Joker +0.21 points 20,004 observations 19,996 observations
Shed High -0.13 points 19,987 observations 20,013 observations
Shed Tiles +0.07 points 19,998 observations 20,002 observations
Shed Points -0.04 points 20,001 observations 19,999 observations
Shape Rack +0.00 points 19,996 observations 20,004 observations

Limits and source files

  • The confirmation applies to Classic Rummikub, four research-strength agents, and their fixed all-distinct table.
  • No preregistered finalist comparison reached the three-percentage-point practical threshold or passed its adjusted statistical test.
  • Greedy tile shedding had the highest descriptive win rate in confirmation, but that specific comparison was not preregistered and remains a lead for future study.
  • The policies take exact rack-clearing wins and optimize their declared rack objective. They do not solve every long-range consequence of a table arrangement.
  • The broad screen and fixed finalist table answer different questions. Strategy strength can change with the mix of opponents.
Exploration
b21341d8cd1af627870111d9c28fc38e718dfbc615d07f9bc512e2ee99cf1d32
Confirmation
b8ca86c326b0d880c2e64d4e0933b87be5de0c984346b47c4c7226c726792e12
Tactical gate
106c498c764d1216d1097663311013c53a919b9d97d1d3838c3b8f40e9330ade
Rule profile
classic_v1