0ac69b101a0a4515dc828715f2da24266eb2d200
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0ac69b101a | Replace bot tuning guesses with enforced calibration | ||
|
|
fffd60ad9d |
Tune the skill axis: profiles now play like their labels
The reported symptom was a "Rock" at 41% VPIP against a 12% setting. Chasing it uncovered four separate places where a *skill* parameter was smuggling in a *style* change — the two axes were supposed to be independent. 1. Error direction. blunder() pushed every mistake the same way, so a 30% error rate put any beginner near a 30% VPIP floor regardless of style. Making it style-directed fixed the Rock but collapsed the gradient, because a tight player's errors then became folds, which cost almost nothing. Errors are now split: pre-flop follows the player's character (a nit's mistake is folding a hand they should have played), while post-flop stays costly for everyone — paying off when beaten and checking back hands worth betting. That is also where weak players genuinely lose money. 2. potOddsRespect shifted weak players systematically toward calling. That is not a weakness — calling wider than break-even against bad opponents is a winning adjustment, so it handed low-skill bots a real edge. Discipline now means ACCURACY: a weak player misjudges the threshold in either direction. 3. OpponentModel treated 0.5 as a neutral bet/raise share. Folds, checks and calls are counted too, so a normal player sits near 0.32 — every opponent read as passive, and the only two levels that consult the model tightened against the whole table and lost money for it. The exploitation feature was a handicap. Baseline calibrated and named. 4. positionAwareness widened 45% in position but narrowed 25% out of it. A seat is last to act about a quarter of the time, so the tighter branch dominated and higher awareness silently meant fewer hands. Position now shifts WHERE hands are played, not how many. Also: the pre-flop slop multiplier now saturates (multiplying pushed a 0.75 maniac to 0.93 while still drowning out the tight end), raw pot odds carry an implied-odds discount, and skill levels are re-spaced. Results: Rock 41.3% -> 14.5% VPIP, every style ordered correctly by looseness, and win rates down from ~113 to ~20 bb/100 for a strong seat. HONEST LIMITATION: Advanced and Expert are not separable. Over 100k hands their order flips with the seed. The simulator now asserts each level beats the one two tiers below it — true on every seed tried — rather than strict adjacent ordering, which would be reading noise as signal. Separating the top two needs either a wider parameter gap or a different distinguishing mechanism. New ProfileBehaviourTest is the regression that was missing: it asserts styles actually produce their own behaviour. A gradient can look healthy while every profile is misnamed, which is exactly what happened. Tests: 154 -> 166 (75 engine JVM, 75 Android host, 16 app). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dc136b7ecc |
Route every call-resolving path through canonicalCall
Direct calls were canonicalised last commit, but the two downgrade paths inside the BET/RAISE branch still built Action(CALL, toCall) with the unclamped amount: an illegal raise falling back to a call, and a player barred from re-raising after an incomplete all-in. A short stack there recorded chips that never moved. That path matters more than the direct one — it is where malformed agent output lands, which is precisely what an LLM will produce once the persona layer exists. All three now go through one canonicalCall(seat, toCall). There is exactly one construction of a CALL action left in the file, which is the point: three copies of this rule is how it went wrong twice. Regressions for both paths, each verified to fail with only its own path reverted — recording 90 instead of 30, and 30 instead of 20. Tests: 150 -> 154 (69 engine JVM, 69 Android host, 16 app). Simulation figures unchanged (289.12 / 113.19 / -195.64), chips conserved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0bf57bf8d |
One clamp, one contract: callAmount owns what a call costs
The visible label was right but the contract behind it was not. The clamp lived in three places — DecisionOffer.callAmount (unclamped, so it still reported the full amount owed), buttonsFor(), and PokerViewModel — and only the presentation copy was tested. Three copies of a rule is how the original mismatch happened. DecisionOffer.callAmount is now minOf(toCall, stack), with callIsAllIn beside it, and both the label and the submitted action read from it. Nothing recomputes the clamp. Engine sanitisation also left an oversized CALL amount unchanged in HandEvent even though commit() caps the commitment, so history described chips that never moved. Since replay and the coach both read that history, calls are now recorded at what was actually committed. Note the arithmetic: heads-up, seat 1 is the big blind, so with a 40 stack facing a raise to 100 the call commits the remaining 30, not 40. My first test asserted 40 and the engine was right. Verified by reverting: both contract fixes fail their tests. Tests: 144 -> 150 (67 engine JVM, 67 Android host, 16 app). Simulation figures unchanged (289.12 / 113.19 / -195.64), chips conserved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4f1ef45dd7 |
Quote the real price when a call exceeds the stack
Facing more than you have left is a call for your whole stack — commit() clamps the commitment — so labelling it "Call 35" when only 20 can be paid promised a price the player could neither make nor owed. It now reads "All in 20". The ViewModel also submits the clamped amount rather than relying on the engine to fix it up, so the action sent never disagrees with the label shown. Verified by reverting: both new label tests fail. Tests: 141 -> 144 (64 engine JVM, 64 Android host, 16 app). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f0c8a7c430 |
Fix legal-raise omission, hand header, and decision identity
1. Raise could vanish when it was legal. The action bar required maxRaiseTo > minRaiseTo, which hid an exact-minimum raise and any legal short all-in — both of which the engine accepts (sizeBet clamps a stack too short for a full min-raise to maxRaiseTo). The button now shows whenever the engine would accept a raise; the slider only appears when there is a genuine range, and a single-size raise submits a fixed amount. 2. The header used handsPlayed + 1, which increments when a hand *finishes*, so during the showdown hold it labelled hand 1's result as "Hand 2". It now uses snapshot.handNumber, which is by construction the hand being displayed. 3. Decision identity is now owned by the engine. HumanAgent minted its own tokens, so liveOffer() had to approximate matching with hand + street + seat — ambiguous, because a player can face two decisions on one street (bet, get raised, act again). Table stamps one monotonic token per decision, carried on both DecisionContext and TableSnapshot.toActToken, so liveOffer() matches exactly and the residual race is gone rather than narrowed. 4. Presentation logic moved out of the composable into buttonsFor(), a pure function of the offer, so it is testable without a Compose runtime. The app module had no tests at all; it now has 13 covering raise visibility, fold suppression, labels, and every liveOffer() matching case. Each new test was verified to fail with its fix reverted: reverting raise visibility and token matching failed exactly four, and no others. Tests: 62 -> 141 total (64 engine JVM, 64 Android host, 13 app). Simulation unchanged, chips conserved. Verified on the emulator; physical device untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7382bc8638 |
Fix the fold experience: honest actions, matched state, action identity
Reported as "folding looks broken". It was five separate defects.
1. Fold silently became Check. sanitise() rewrote FOLD to CHECK whenever
checking was free, so a UI showing a Fold button folded nothing and the
player kept being asked to act. Folding is legal at any turn — it simply
mucks — so FOLD is now honoured literally. The engine must never substitute a
different action than the caller asked for. The reverse rewrite (an illegal
CHECK facing a bet becoming FOLD) is legitimate and stays.
Safe by construction: no bot emits FOLD when it can check, and the 20k-hand
simulation reproduces byte-identical numbers (289.12 / 113.19 / -195.64).
2. Snapshot and offer could describe different moments. The 32-deep frame
channel let the engine race far ahead of the animation, so the action on
offer could belong to a later street, or another hand. The channel is now
RENDEZVOUS, capping the engine at one frame ahead, and UiState.liveOffer()
only surfaces an offer whose hand and street match the table on screen.
3. Stale and double taps could act on a later decision. DecisionOffer now
carries a token; submit() requires it and rejects anything stale, so a second
tap is dropped rather than applied to whatever comes next.
4. A real fold was invisible. The hero kept normal cards and no folded state, so
a correctly processed fold looked like a bug. Cards now dim, FOLDED shows in
red, and the action bar explains the player is sitting out.
5. Non-atomic UiState updates from two coroutines now use update {}.
Also: Fold is hidden when checking is free (folding a free hand is never
correct, and offering it invites an accidental muck), and onCleared no longer
calls human.cancel() — viewModelScope is already cancelled by then so the launch
never ran; scope cancellation already propagates into act()'s finally.
The delivery tests were weak as charged: no slow consumer, and not the app's
capacity. Replaced with a genuinely slow consumer measuring how far the engine
runs ahead — asserting <= 1 on RENDEZVOUS, and > 1 on a 32-deep buffer to
document why the buffer was removed.
Verified on the emulator (physical device untouched): folded facing a bet, hero
showed FOLDED, was never asked again that hand, and play advanced to hand 2.
Tests: 55 -> 62, green on jvmTest and testAndroidHostTest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
950c7ceb57 |
Playable Android table
The game runs on device: verified on a Pixel 10 Pro emulator (Android 17) by
installing, tapping through a hand, and confirming it advanced pre-flop to flop
with correct pot, folds, and re-offered action.
App:
- :app module on AGP 9.2.1. Note AGP 9 has built-in Kotlin support, so applying
org.jetbrains.kotlin.android conflicts with it ("extension with name 'kotlin'
already registered"); only android.application + kotlin.compose are applied,
matching recipeze.
- PokerViewModel runs a continuous cash game and publishes to Compose.
- Compose table: opponents, board, pot, hero, action bar with a raise slider.
Frames are queued, not conflated. An all-in runout emits flop, turn and river
microseconds apart; pushing those into a StateFlow would collapse them and the
board would jump from empty to complete. The engine's suspending observer sends
into a Channel, a consumer paces each frame, and only then is StateFlow updated
— so backpressure paces the engine rather than the UI dropping frames. Three
tests cover this, including a characterisation test showing a conflating
StateFlow does lose the intermediate frames.
Assets:
- tools/generate_card_assets.sh rasterises the SVGs into four density buckets
using sips, which renders SVG directly — no librsvg or ImageMagick.
- Resource names are prefixed card_ because Android resource names may not start
with a digit (10_of_clubs would be rejected).
- CardArt.kt maps deck index to drawable via static R references, so R8 resource
shrinking cannot strip the artwork the way getIdentifier lookups would risk.
Layout fixes found by actually looking at the running app: five opponents did
not fit a fixed-width scrolling row (Enzo was off-screen), the header collided
with the status bar clock, and the board floated against a large dead space.
Tests: 52 -> 55, green on jvmTest and testAndroidHostTest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
679b25e2c7 |
Publish the all-in runout street by street
During an all-in runout, dealRemainingBoard() dealt every remaining card without touching currentStreet, so a pre-flop all-in produced a terminal snapshot labelled PREFLOP carrying a five-card board. Setting currentStreet = RIVER would fix the label but leave a second problem: the board jumped from empty to complete in a single snapshot, so the UI could not animate the runout — the moment a poker table most needs to. Instead dealRemainingBoard() is now suspend and publishes each street as it lands, which keeps currentStreet honest as a consequence rather than as a special case. Each runout street also sweeps its betting into the pot via prepareRound(), matching the normal street transition. Verified by reverting: the board went 0 -> 5 with the terminal snapshot labelled PREFLOP, and three of the four new tests failed. Tests: 48 -> 52, green on jvmTest and testAndroidHostTest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
eefcd5966c |
Fix UI-boundary defects; run tests on the Android variant too
Two defects found reviewing the UI boundary before Compose work: 1. HumanAgent.awaiting was a plain mutable property written by the game coroutine and read by the UI — a data race, and invisible to Compose. It also exposed DecisionContext, which holds a live Seat whose fields mutate as the hand proceeds, so even a safe read could observe torn state. Replaced with an immutable DecisionOffer published through a StateFlow. A test mutates the live seat after publication and asserts the offer does not change. 2. STREET_COMPLETE was emitted after dealing the new street but before the round state was reset, so a flop snapshot carried pre-flop currentBet and committedThisRound — the UI would have painted last street's chips in front of every player alongside the new board. The reset is now prepareRound(), called before publishing. Verified: with the ordering reverted the new test fails with currentBet 10 on the flop. Also: - HumanAgent.cancel() is now covered directly; the previous test only cancelled the coroutine running act(). cancel() and submit() both report whether anything was actually pending. - Android host tests enabled via withHostTestBuilder, so the shared suite runs against the Android variant instead of the AAR merely compiling. No librsvg needed for card assets: sips rasterizes the SVGs directly at exact 2:3 dimensions, court cards and patterned backs included. Tests: 45 -> 48, now green on both jvmTest and testAndroidHostTest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7f82251d86 |
Android target, suspending agents, and observable table snapshots
Build: - :engine now uses com.android.kotlin.multiplatform.library (AGP 9.2.1), the modern KMP Android integration, rather than plain androidTarget(). Produces engine.aar alongside the JVM target; compileAndroidMain verified. - Version catalog added; SDK levels match the other JSJ apps (compileSdk 37, minSdk 26). Engine: - PlayerAgent.act() and Table.playHand() are now suspend, so a human player can wait for input without blocking a thread. Bots are unaffected; the simulator wraps in runBlocking. - TableSnapshot/SeatSnapshot published after the deal, before and after every action, and at the finish. Immutable, aliasing no live Seat state, giving animation, hand history, saving, and replay one boundary to work against. - Snapshots carry the whole truth; maskedFor(viewer) is an explicit step that hides hole cards the viewer is not entitled to. Showdown reveals contenders; folded hands never are. - Terminal snapshots report the contested pot rather than 0. settle() zeroes contributions when awarding, so the naive value was empty at exactly the moment the UI needs to show what was won. Caught by a new test. - HumanAgent suspends on a CompletableDeferred and clears its pending state in a finally block, so cancelling an abandoned hand releases the wait instead of stranding it. Re-usable afterwards; a stale submit returns false. Tests: 37 -> 45. Chips still conserved, skill gradient still monotonic. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
96af198097 |
Correct PreflopChart percentile docs
The KDoc claimed the worst hand is percentile 1.0, but percentiles mark the START of each class's band, so the last class begins at (1326-12)/1326 ~ 0.991 and no real hand reaches 1.0. Documents that this is correct for gating, since a class is admitted when its band opens inside the range. Also notes that the unused key slots keep 1.0 as an unreachable sentinel. Docs only; no behaviour change. 37 tests still pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1222401fb |
Fix four rules and modelling defects found in review
TDA Rule 47 — cumulative incomplete raises: A boolean could not express "facing at least a full raise since acting", so several short all-ins that together reached a full raise failed to reopen betting. Seat now records lastActedAtBet (the currentBet when it last acted); mayRaise() reopens when currentBet - lastActedAtBet >= minRaiseSize. This subsumes the single-incomplete-raise case, so the hasActed reset in apply() is gone. TDA Rule 20 — odd chips: Split-pot remainders were awarded in seat-list order. They now go to the first winning seat clockwise from the button. The old test also never produced an odd pot (20 chips heads-up), so it only ever proved an even split; it now builds a genuinely odd 25-chip pot via a folded small blind and asserts which seat takes the extra chip. OpponentModel skipped events across hands: It inferred a new hand from a shrinking history, but history is cleared each hand: having consumed 3 events, first observing the next hand at 4 events left 4 < 3 false and silently dropped the first three. observe() now takes an explicit handNumber, exposed via DecisionContext and Table.handNumber. PreflopChart percentile semantics: The 169 classes were ranked equally, but they are not equally likely — a pair is 6 of 1326 combinations, suited 4, offsuit 12. "Top 12%" therefore meant 12% of classes, not of dealt hands, so looseness did not mean what it claimed. Percentiles are now weighted by combination count. Tests: 30 -> 37. Each new test was verified to fail with its fix reverted. Skill gradient still monotonic: 85.9 / 79.8 / 17.4 / -183.1 bb/100 over 50k hands, chips conserved on both tables. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
479be1f6b9 |
Initial commit: Hold'em engine, bots, and simulation harness
Kotlin Multiplatform engine (JVM target only for now; androidTarget and iosArm64 slot in without touching commonMain). Core: - HandEvaluator: single-pass 5-7 card evaluation, ~24M evals/sec. Verified exhaustively against published frequencies for all 2,598,960 five-card hands. - Equity: Monte Carlo with ties split. PreflopChart ranks the 169 starting hands using all-in equity plus an explicit playability adjustment, so looseness means "plays the top N%". - Table: no-limit betting rounds, side pots, odd-chip splits, uncalled-bet refunds, and incomplete (short all-in) raises that correctly do not reopen betting. Bots: - SkillLevel and PlayStyle are orthogonal axes. Skill drives decision quality (rollout accuracy, pot-odds discipline, position awareness, error rate); style drives bluffing, sandbagging, aggression, tightness. - BotMood gives tilt that persists between hands and decays. - OpponentModel lets Advanced/Expert exploit habitual bettors. - MathBot emits a DecisionTrace of the numbers behind each decision, which the coach will later hand to an LLM to narrate. The LLM never does poker maths. Simulator: - 2,200-3,400 hands/sec. Deck RNG is separate from bot RNGs so rollout counts cannot shift the deal. - Controlled skill-ladder test asserts the difficulty gradient is monotonic: 73.9 / 53.9 / 27.6 / -155.4 bb/100 over 50k hands. Assets: 52 CC0 English-pattern card faces plus generated backs. Tests: 30 passing (evaluator, table rules, pre-flop chart). Known open: win-rate magnitudes ~10x realistic and several profiles looser than their labels. Tuning, not correctness. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |