fffd60ad9d0e266e9f954bac9acbad47e477a848
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fffd60ad9d |
Tune the skill axis: profiles now play like their labels
The reported symptom was a "Rock" at 41% VPIP against a 12% setting. Chasing it uncovered four separate places where a *skill* parameter was smuggling in a *style* change — the two axes were supposed to be independent. 1. Error direction. blunder() pushed every mistake the same way, so a 30% error rate put any beginner near a 30% VPIP floor regardless of style. Making it style-directed fixed the Rock but collapsed the gradient, because a tight player's errors then became folds, which cost almost nothing. Errors are now split: pre-flop follows the player's character (a nit's mistake is folding a hand they should have played), while post-flop stays costly for everyone — paying off when beaten and checking back hands worth betting. That is also where weak players genuinely lose money. 2. potOddsRespect shifted weak players systematically toward calling. That is not a weakness — calling wider than break-even against bad opponents is a winning adjustment, so it handed low-skill bots a real edge. Discipline now means ACCURACY: a weak player misjudges the threshold in either direction. 3. OpponentModel treated 0.5 as a neutral bet/raise share. Folds, checks and calls are counted too, so a normal player sits near 0.32 — every opponent read as passive, and the only two levels that consult the model tightened against the whole table and lost money for it. The exploitation feature was a handicap. Baseline calibrated and named. 4. positionAwareness widened 45% in position but narrowed 25% out of it. A seat is last to act about a quarter of the time, so the tighter branch dominated and higher awareness silently meant fewer hands. Position now shifts WHERE hands are played, not how many. Also: the pre-flop slop multiplier now saturates (multiplying pushed a 0.75 maniac to 0.93 while still drowning out the tight end), raw pot odds carry an implied-odds discount, and skill levels are re-spaced. Results: Rock 41.3% -> 14.5% VPIP, every style ordered correctly by looseness, and win rates down from ~113 to ~20 bb/100 for a strong seat. HONEST LIMITATION: Advanced and Expert are not separable. Over 100k hands their order flips with the seed. The simulator now asserts each level beats the one two tiers below it — true on every seed tried — rather than strict adjacent ordering, which would be reading noise as signal. Separating the top two needs either a wider parameter gap or a different distinguishing mechanism. New ProfileBehaviourTest is the regression that was missing: it asserts styles actually produce their own behaviour. A gradient can look healthy while every profile is misnamed, which is exactly what happened. Tests: 154 -> 166 (75 engine JVM, 75 Android host, 16 app). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
950c7ceb57 |
Playable Android table
The game runs on device: verified on a Pixel 10 Pro emulator (Android 17) by
installing, tapping through a hand, and confirming it advanced pre-flop to flop
with correct pot, folds, and re-offered action.
App:
- :app module on AGP 9.2.1. Note AGP 9 has built-in Kotlin support, so applying
org.jetbrains.kotlin.android conflicts with it ("extension with name 'kotlin'
already registered"); only android.application + kotlin.compose are applied,
matching recipeze.
- PokerViewModel runs a continuous cash game and publishes to Compose.
- Compose table: opponents, board, pot, hero, action bar with a raise slider.
Frames are queued, not conflated. An all-in runout emits flop, turn and river
microseconds apart; pushing those into a StateFlow would collapse them and the
board would jump from empty to complete. The engine's suspending observer sends
into a Channel, a consumer paces each frame, and only then is StateFlow updated
— so backpressure paces the engine rather than the UI dropping frames. Three
tests cover this, including a characterisation test showing a conflating
StateFlow does lose the intermediate frames.
Assets:
- tools/generate_card_assets.sh rasterises the SVGs into four density buckets
using sips, which renders SVG directly — no librsvg or ImageMagick.
- Resource names are prefixed card_ because Android resource names may not start
with a digit (10_of_clubs would be rejected).
- CardArt.kt maps deck index to drawable via static R references, so R8 resource
shrinking cannot strip the artwork the way getIdentifier lookups would risk.
Layout fixes found by actually looking at the running app: five opponents did
not fit a fixed-width scrolling row (Enzo was off-screen), the header collided
with the status bar clock, and the board floated against a large dead space.
Tests: 52 -> 55, green on jvmTest and testAndroidHostTest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f1222401fb |
Fix four rules and modelling defects found in review
TDA Rule 47 — cumulative incomplete raises: A boolean could not express "facing at least a full raise since acting", so several short all-ins that together reached a full raise failed to reopen betting. Seat now records lastActedAtBet (the currentBet when it last acted); mayRaise() reopens when currentBet - lastActedAtBet >= minRaiseSize. This subsumes the single-incomplete-raise case, so the hasActed reset in apply() is gone. TDA Rule 20 — odd chips: Split-pot remainders were awarded in seat-list order. They now go to the first winning seat clockwise from the button. The old test also never produced an odd pot (20 chips heads-up), so it only ever proved an even split; it now builds a genuinely odd 25-chip pot via a folded small blind and asserts which seat takes the extra chip. OpponentModel skipped events across hands: It inferred a new hand from a shrinking history, but history is cleared each hand: having consumed 3 events, first observing the next hand at 4 events left 4 < 3 false and silently dropped the first three. observe() now takes an explicit handNumber, exposed via DecisionContext and Table.handNumber. PreflopChart percentile semantics: The 169 classes were ranked equally, but they are not equally likely — a pair is 6 of 1326 combinations, suited 4, offsuit 12. "Top 12%" therefore meant 12% of classes, not of dealt hands, so looseness did not mean what it claimed. Percentiles are now weighted by combination count. Tests: 30 -> 37. Each new test was verified to fail with its fix reverted. Skill gradient still monotonic: 85.9 / 79.8 / 17.4 / -183.1 bb/100 over 50k hands, chips conserved on both tables. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
479be1f6b9 |
Initial commit: Hold'em engine, bots, and simulation harness
Kotlin Multiplatform engine (JVM target only for now; androidTarget and iosArm64 slot in without touching commonMain). Core: - HandEvaluator: single-pass 5-7 card evaluation, ~24M evals/sec. Verified exhaustively against published frequencies for all 2,598,960 five-card hands. - Equity: Monte Carlo with ties split. PreflopChart ranks the 169 starting hands using all-in equity plus an explicit playability adjustment, so looseness means "plays the top N%". - Table: no-limit betting rounds, side pots, odd-chip splits, uncalled-bet refunds, and incomplete (short all-in) raises that correctly do not reopen betting. Bots: - SkillLevel and PlayStyle are orthogonal axes. Skill drives decision quality (rollout accuracy, pot-odds discipline, position awareness, error rate); style drives bluffing, sandbagging, aggression, tightness. - BotMood gives tilt that persists between hands and decays. - OpponentModel lets Advanced/Expert exploit habitual bettors. - MathBot emits a DecisionTrace of the numbers behind each decision, which the coach will later hand to an LLM to narrate. The LLM never does poker maths. Simulator: - 2,200-3,400 hands/sec. Deck RNG is separate from bot RNGs so rollout counts cannot shift the deal. - Controlled skill-ladder test asserts the difficulty gradient is monotonic: 73.9 / 53.9 / 27.6 / -155.4 bb/100 over 50k hands. Assets: 52 CC0 English-pattern card faces plus generated backs. Tests: 30 passing (evaluator, table rules, pre-flop chart). Known open: win-rate magnitudes ~10x realistic and several profiles looser than their labels. Tuning, not correctness. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |