Completed exploration and held-out confirmation
The dramatic early result did not survive the final test.
This page keeps the exploratory ranking, the separate finalist test, and the tactical audit visible. The distinction matters: the 100,000-game screen found possibilities; the untouched 24,000-game confirmation decided which claims could be published.
- Total games
- 124,000
- Exploration
- 100,000
- Confirmation
- 24,000
- Fresh deal clusters
- 1,000
Confirmed answer
No preregistered finalist comparison showed a meaningful advantage.
The four finalists ranged from 24.6% to 25.7% in descriptive win rate. Every preregistered difference included zero in its 95% clustered interval, and none reached the study’s three-percentage-point threshold.
Table Denial’s 8.11-point exploratory lead over Balanced became a 0.34-point deficit in confirmation. The likely lesson is context: a reactive strategy can thrive in one opponent ecology and look ordinary in another.
Start here
Three kinds of evidence
- Tactical audit
- Checks whether the engine and serious policies miss an immediate way to empty the rack. They missed none in 10,000 test positions.
- Exploration
- Tries many strategies and table mixes to find leads. Its rankings are useful for choosing the next test, not for declaring a winner.
- Confirmation
- Tests frozen questions on fresh deals. Only this stage can promote the preregistered comparisons to findings.
- Percentage points
- A direct win-rate change. Moving from 25% to 28% is a gain of three percentage points.
Not sure what a strategy name means? Read the plain-English strategy field guide first.
Held-out result
The four finalists finished close together
These are descriptive rates from the same cyclically balanced finalist table. Overlapping intervals mean the visible order should not be read as a proven ranking.
| Strategy | Win rate | 95% interval | Mean score | Losing rack |
|---|---|---|---|---|
| Greedy Tiles | 25.7% | 25.1% to 26.2% | +0.42 | 28.30 |
| Balanced | 25.0% | 24.5% to 25.6% | +0.37 | 27.68 |
| Table Denial | 24.7% | 24.2% to 25.2% | -0.23 | 27.98 |
| Opponent Modeler | 24.6% | 24.0% to 25.1% | -0.56 | 28.17 |
Preregistered comparisons
What the final test actually asked
Each interval is clustered by the 1,000 independent deal blocks. A claim needed both an adjusted statistical result and an advantage of at least three percentage points.
| Comparison | Difference | 95% interval | Result |
|---|---|---|---|
| Table Denial vs Balanced | -0.34 points | -1.06 to +0.39 | Not supported |
| Table Denial vs Opponent Modeler | +0.12 points | -0.35 to +0.58 | Not supported |
| Opponent Modeler vs Balanced | -0.45 points | -1.17 to +0.27 | Not supported |
Exploratory screen
Why Table Denial reached the final
The broad screen mixed homogeneous, two-versus-two, three-plus-one, and all-distinct tables. These results generated hypotheses. They do not override the held-out result above.
| Strategy | Win rate | Mean score | Rack value left |
|---|---|---|---|
| Table Denial | 33.5% | +13.59 | 17.21 |
| Opponent Modeler | 28.3% | +6.11 | 19.17 |
| Match Aware | 26.9% | +5.83 | 19.51 |
| Strategic Drawer | 26.7% | -2.33 | 26.42 |
| Manipulation Averse | 26.6% | -8.75 | 30.14 |
| Greedy Tiles | 26.1% | +4.51 | 20.54 |
| Shallow Planner | 25.8% | +4.82 | 18.97 |
| Endgame Liquidator | 25.8% | +2.13 | 20.47 |
| Greedy Points | 25.8% | +5.09 | 19.01 |
| Eager Beginner | 25.7% | +4.01 | 20.50 |
| Conservative | 25.4% | +3.78 | 19.89 |
| Balanced | 25.4% | +3.60 | 19.45 |
| Late Opening Hoarder | 25.1% | -7.89 | 30.67 |
| Joker Holder | 25.0% | +3.64 | 18.93 |
| Pivot Copy Keeper | 24.9% | -1.31 | 23.57 |
| Joker Farmer | 24.7% | +0.28 | 20.60 |
| Random Legal | 4.4% | -35.40 | 39.50 |
Controlled ingredients
No single preference moved wins by even one point
A separate 32-combination experiment switched five tendencies on and off. All observed differences were smaller than 0.21 percentage points, another warning against one-rule strategy slogans.
| Tendency emphasized | Win-rate difference | High setting | Low setting |
|---|---|---|---|
| Hold Joker | +0.21 points | 20,004 observations | 19,996 observations |
| Shed High | -0.13 points | 19,987 observations | 20,013 observations |
| Shed Tiles | +0.07 points | 19,998 observations | 20,002 observations |
| Shed Points | -0.04 points | 20,001 observations | 19,999 observations |
| Shape Rack | +0.00 points | 19,996 observations | 20,004 observations |
Read before quoting
Limits and source files
- The confirmation applies to Classic Rummikub, four research-strength agents, and their fixed all-distinct table.
- No preregistered finalist comparison reached the three-percentage-point practical threshold or passed its adjusted statistical test.
- Greedy tile shedding had the highest descriptive win rate in confirmation, but that specific comparison was not preregistered and remains a lead for future study.
- The policies take exact rack-clearing wins and optimize their declared rack objective. They do not solve every long-range consequence of a table arrangement.
- The broad screen and fixed finalist table answer different questions. Strategy strength can change with the mix of opponents.
- Exploration
b21341d8cd1af627870111d9c28fc38e718dfbc615d07f9bc512e2ee99cf1d32- Confirmation
b8ca86c326b0d880c2e64d4e0933b87be5de0c984346b47c4c7226c726792e12- Tactical gate
106c498c764d1216d1097663311013c53a919b9d97d1d3838c3b8f40e9330ade- Rule profile
classic_v1