Tune the skill axis: profiles now play like their labels
The reported symptom was a "Rock" at 41% VPIP against a 12% setting. Chasing it uncovered four separate places where a *skill* parameter was smuggling in a *style* change — the two axes were supposed to be independent. 1. Error direction. blunder() pushed every mistake the same way, so a 30% error rate put any beginner near a 30% VPIP floor regardless of style. Making it style-directed fixed the Rock but collapsed the gradient, because a tight player's errors then became folds, which cost almost nothing. Errors are now split: pre-flop follows the player's character (a nit's mistake is folding a hand they should have played), while post-flop stays costly for everyone — paying off when beaten and checking back hands worth betting. That is also where weak players genuinely lose money. 2. potOddsRespect shifted weak players systematically toward calling. That is not a weakness — calling wider than break-even against bad opponents is a winning adjustment, so it handed low-skill bots a real edge. Discipline now means ACCURACY: a weak player misjudges the threshold in either direction. 3. OpponentModel treated 0.5 as a neutral bet/raise share. Folds, checks and calls are counted too, so a normal player sits near 0.32 — every opponent read as passive, and the only two levels that consult the model tightened against the whole table and lost money for it. The exploitation feature was a handicap. Baseline calibrated and named. 4. positionAwareness widened 45% in position but narrowed 25% out of it. A seat is last to act about a quarter of the time, so the tighter branch dominated and higher awareness silently meant fewer hands. Position now shifts WHERE hands are played, not how many. Also: the pre-flop slop multiplier now saturates (multiplying pushed a 0.75 maniac to 0.93 while still drowning out the tight end), raw pot odds carry an implied-odds discount, and skill levels are re-spaced. Results: Rock 41.3% -> 14.5% VPIP, every style ordered correctly by looseness, and win rates down from ~113 to ~20 bb/100 for a strong seat. HONEST LIMITATION: Advanced and Expert are not separable. Over 100k hands their order flips with the seed. The simulator now asserts each level beats the one two tiers below it — true on every seed tried — rather than strict adjacent ordering, which would be reading noise as signal. Separating the top two needs either a wider parameter gap or a different distinguishing mechanism. New ProfileBehaviourTest is the regression that was missing: it asserts styles actually produce their own behaviour. A gradient can look healthy while every profile is misnamed, which is exactly what happened. Tests: 154 -> 166 (75 engine JVM, 75 Android host, 16 app). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -42,8 +42,14 @@ export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
|
||||
3. **Pre-flop is range-based, not equity-based.** All-in equity overvalues trash
|
||||
(7-2o has ~35% vs one random hand but is unplayable). `PreflopChart` ranks the
|
||||
169 starting hands so `looseness` means "plays the top N%".
|
||||
4. **The simulator is how bots get tuned.** Run it after any bot change; it prints
|
||||
a controlled skill-ladder test that must stay monotonic.
|
||||
4. **The simulator is how bots get tuned.** Run it after any bot change, across
|
||||
several seeds — a single seed will happily agree with a wrong conclusion.
|
||||
5. **A skill parameter must not smuggle in a style change.** Several bugs came
|
||||
from exactly this: `positionAwareness` silently reduced hands played,
|
||||
`potOddsRespect` systematically loosened weak players (which is a *winning*
|
||||
adjustment, so it inverted the gradient), and error direction overwrote style
|
||||
entirely. Skill should change how *well* a decision is made, not how loose or
|
||||
tight the player is.
|
||||
|
||||
## Testing notes
|
||||
|
||||
@@ -64,8 +70,14 @@ export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
|
||||
action order, malformed agent output, **TDA Rule 47** (incomplete raises do not
|
||||
reopen betting, but several that cumulatively reach a full raise do), and
|
||||
**TDA Rule 20** (odd chip to the first winner left of the button).
|
||||
- Bots: skill gradient **passes** monotonically (85.9 / 79.8 / 17.4 / −183.1
|
||||
bb/100 at 50k hands).
|
||||
- Bots: profiles play like their labels (Rock 14.5% VPIP against a 12% setting;
|
||||
`ProfileBehaviourTest` asserts this). Win rates are in a plausible range —
|
||||
roughly +20 bb/100 for a strong seat rather than the earlier +113.
|
||||
- **Adjacent top tiers are not separable.** Advanced and Expert sit inside
|
||||
seed-to-seed noise of each other over 100k hands. The simulator therefore
|
||||
asserts each level beats the one *two* tiers below it, which holds on every
|
||||
seed tried; claiming strict adjacent ordering from one seed would be reading
|
||||
noise as signal.
|
||||
|
||||
### Rules invariants that are easy to get wrong
|
||||
|
||||
|
||||
Reference in New Issue
Block a user