Files
holdem_poker/CLAUDE.md
T
thejayman77 f1222401fb Fix four rules and modelling defects found in review
TDA Rule 47 — cumulative incomplete raises:
A boolean could not express "facing at least a full raise since acting", so
several short all-ins that together reached a full raise failed to reopen
betting. Seat now records lastActedAtBet (the currentBet when it last acted);
mayRaise() reopens when currentBet - lastActedAtBet >= minRaiseSize. This
subsumes the single-incomplete-raise case, so the hasActed reset in apply() is
gone.

TDA Rule 20 — odd chips:
Split-pot remainders were awarded in seat-list order. They now go to the first
winning seat clockwise from the button. The old test also never produced an odd
pot (20 chips heads-up), so it only ever proved an even split; it now builds a
genuinely odd 25-chip pot via a folded small blind and asserts which seat takes
the extra chip.

OpponentModel skipped events across hands:
It inferred a new hand from a shrinking history, but history is cleared each
hand: having consumed 3 events, first observing the next hand at 4 events left
4 < 3 false and silently dropped the first three. observe() now takes an
explicit handNumber, exposed via DecisionContext and Table.handNumber.

PreflopChart percentile semantics:
The 169 classes were ranked equally, but they are not equally likely — a pair is
6 of 1326 combinations, suited 4, offsuit 12. "Top 12%" therefore meant 12% of
classes, not of dealt hands, so looseness did not mean what it claimed.
Percentiles are now weighted by combination count.

Tests: 30 -> 37. Each new test was verified to fail with its fix reverted.
Skill gradient still monotonic: 85.9 / 79.8 / 17.4 / -183.1 bb/100 over 50k
hands, chips conserved on both tables.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 05:03:11 -04:00

3.9 KiB
Raw Blame History

Poker — Texas Hold'em with teachable AI opponents

Kotlin Multiplatform. Ships iOS + Android; Android first (only Android hardware for physical testing).

Build

No java/gradle on PATH — use Android Studio's bundled JDK:

export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
./gradlew :engine:jvmTest          # evaluator + engine tests
./gradlew :sim:run --args="50000"  # simulate 50k hands, print bot stats

Layout

Path What
engine/src/commonMain/.../core/ Cards, evaluator, equity, pre-flop chart
engine/src/commonMain/.../bot/ Skill/style profiles, MathBot
engine/src/commonMain/.../game/ Table — betting rounds, side pots, showdown
sim/ JVM-only headless simulator used to tune bot profiles
assets/cards/ 52 CC0 card faces + generated backs

engine is pure Kotlin with no platform APIs, so androidTarget() / iosArm64() slot in without touching commonMain.

Design rules

  1. Poker maths never goes near the LLM. Difficulty and style are engine-side EV/frequency calculations — instant, deterministic, testable, offline. The LLM only narrates numbers the engine already computed (DecisionTrace), and adds persona/table talk.
  2. Skill and style are orthogonal. SkillLevel = how correct decisions are; PlayStyle = bluffing, sandbagging, aggression, tightness. Build the strongest bot, then inject controlled error for lower tiers.
  3. Pre-flop is range-based, not equity-based. All-in equity overvalues trash (7-2o has ~35% vs one random hand but is unplayable). PreflopChart ranks the 169 starting hands so looseness means "plays the top N%".
  4. The simulator is how bots get tuned. Run it after any bot change; it prints a controlled skill-ladder test that must stay monotonic.

Testing notes

  • Table takes a CardSource, so StackedDeck.of(holes, board) gives fully deterministic hands. Use it for any rule test.
  • The simulator gives the deck its own RNG, separate from each bot's. Never share one: bots consume RNG proportional to their equityIterations, so a shared stream means changing a profile silently changes the cards dealt.
  • ./gradlew :sim:run --args="chart" dumps the starting-hand ranking.
  • Small samples lie. 1,000 hands is not enough to rank profiles — use 50,000+ before believing a gradient.

Status

  • Evaluator: verified exhaustively against published frequencies for all 2,598,960 five-card hands. ~24M evals/sec.
  • Engine: chip-conserving. Covered by tests: side pots, uncalled-bet refunds, action order, malformed agent output, TDA Rule 47 (incomplete raises do not reopen betting, but several that cumulatively reach a full raise do), and TDA Rule 20 (odd chip to the first winner left of the button).
  • Bots: skill gradient passes monotonically (85.9 / 79.8 / 17.4 / 183.1 bb/100 at 50k hands).

Rules invariants that are easy to get wrong

  • Reopening betting cannot be a boolean. Seat.lastActedAtBet records the bet level a player last acted at; betting reopens when `currentBet - lastActedAtBet

    = minRaiseSize`. Several short all-ins can reach that together.

  • PreflopChart percentiles are weighted by combination counts (pair 6, suited 4, offsuit 12, total 1326), so "top 12%" means 12% of dealt hands, not 12% of the 169 classes.
  • Anything consuming DecisionContext.history across hands must key off handNumber. History is cleared each hand, so a size comparison silently drops events.
  • Known-imperfect: win-rate magnitudes are still ~10x realistic, and several profiles are looser than their labels (the Rock plays ~38% VPIP, should be ~12%). Tuning is the open work.
  • Gradle emits an archives deprecation from the Kotlin Multiplatform plugin's own jvm() target registration — upstream in Kotlin 2.2.10, not our build.
  • Not built yet: LLM persona layer, opt-in coach, Compose UI.