Files
holdem_poker/CLAUDE.md
T
thejayman77 f1222401fb Fix four rules and modelling defects found in review
TDA Rule 47 — cumulative incomplete raises:
A boolean could not express "facing at least a full raise since acting", so
several short all-ins that together reached a full raise failed to reopen
betting. Seat now records lastActedAtBet (the currentBet when it last acted);
mayRaise() reopens when currentBet - lastActedAtBet >= minRaiseSize. This
subsumes the single-incomplete-raise case, so the hasActed reset in apply() is
gone.

TDA Rule 20 — odd chips:
Split-pot remainders were awarded in seat-list order. They now go to the first
winning seat clockwise from the button. The old test also never produced an odd
pot (20 chips heads-up), so it only ever proved an even split; it now builds a
genuinely odd 25-chip pot via a folded small blind and asserts which seat takes
the extra chip.

OpponentModel skipped events across hands:
It inferred a new hand from a shrinking history, but history is cleared each
hand: having consumed 3 events, first observing the next hand at 4 events left
4 < 3 false and silently dropped the first three. observe() now takes an
explicit handNumber, exposed via DecisionContext and Table.handNumber.

PreflopChart percentile semantics:
The 169 classes were ranked equally, but they are not equally likely — a pair is
6 of 1326 combinations, suited 4, offsuit 12. "Top 12%" therefore meant 12% of
classes, not of dealt hands, so looseness did not mean what it claimed.
Percentiles are now weighted by combination count.

Tests: 30 -> 37. Each new test was verified to fail with its fix reverted.
Skill gradient still monotonic: 85.9 / 79.8 / 17.4 / -183.1 bb/100 over 50k
hands, chips conserved on both tables.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 05:03:11 -04:00

83 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Poker — Texas Hold'em with teachable AI opponents
Kotlin Multiplatform. Ships iOS + Android; Android first (only Android hardware
for physical testing).
## Build
No `java`/`gradle` on PATH — use Android Studio's bundled JDK:
```bash
export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
./gradlew :engine:jvmTest # evaluator + engine tests
./gradlew :sim:run --args="50000" # simulate 50k hands, print bot stats
```
## Layout
| Path | What |
|---|---|
| `engine/src/commonMain/.../core/` | Cards, evaluator, equity, pre-flop chart |
| `engine/src/commonMain/.../bot/` | Skill/style profiles, `MathBot` |
| `engine/src/commonMain/.../game/` | `Table` — betting rounds, side pots, showdown |
| `sim/` | JVM-only headless simulator used to **tune** bot profiles |
| `assets/cards/` | 52 CC0 card faces + generated backs |
`engine` is pure Kotlin with no platform APIs, so `androidTarget()` /
`iosArm64()` slot in without touching `commonMain`.
## Design rules
1. **Poker maths never goes near the LLM.** Difficulty and style are engine-side
EV/frequency calculations — instant, deterministic, testable, offline. The LLM
only narrates numbers the engine already computed (`DecisionTrace`), and adds
persona/table talk.
2. **Skill and style are orthogonal.** `SkillLevel` = how correct decisions are;
`PlayStyle` = bluffing, sandbagging, aggression, tightness. Build the strongest
bot, then inject *controlled error* for lower tiers.
3. **Pre-flop is range-based, not equity-based.** All-in equity overvalues trash
(7-2o has ~35% vs one random hand but is unplayable). `PreflopChart` ranks the
169 starting hands so `looseness` means "plays the top N%".
4. **The simulator is how bots get tuned.** Run it after any bot change; it prints
a controlled skill-ladder test that must stay monotonic.
## Testing notes
- `Table` takes a `CardSource`, so `StackedDeck.of(holes, board)` gives fully
deterministic hands. Use it for any rule test.
- The simulator gives the **deck its own RNG**, separate from each bot's. Never
share one: bots consume RNG proportional to their `equityIterations`, so a
shared stream means changing a profile silently changes the cards dealt.
- `./gradlew :sim:run --args="chart"` dumps the starting-hand ranking.
- Small samples lie. 1,000 hands is not enough to rank profiles — use 50,000+
before believing a gradient.
## Status
- Evaluator: verified exhaustively against published frequencies for all
2,598,960 five-card hands. ~24M evals/sec.
- Engine: chip-conserving. Covered by tests: side pots, uncalled-bet refunds,
action order, malformed agent output, **TDA Rule 47** (incomplete raises do not
reopen betting, but several that cumulatively reach a full raise do), and
**TDA Rule 20** (odd chip to the first winner left of the button).
- Bots: skill gradient **passes** monotonically (85.9 / 79.8 / 17.4 / 183.1
bb/100 at 50k hands).
### Rules invariants that are easy to get wrong
- Reopening betting cannot be a boolean. `Seat.lastActedAtBet` records the bet
level a player last acted at; betting reopens when `currentBet - lastActedAtBet
>= minRaiseSize`. Several short all-ins can reach that together.
- `PreflopChart` percentiles are weighted by **combination counts** (pair 6,
suited 4, offsuit 12, total 1326), so "top 12%" means 12% of *dealt hands*, not
12% of the 169 classes.
- Anything consuming `DecisionContext.history` across hands must key off
`handNumber`. History is cleared each hand, so a size comparison silently drops
events.
- Known-imperfect: win-rate magnitudes are still ~10x realistic, and several
profiles are looser than their labels (the Rock plays ~38% VPIP, should be
~12%). Tuning is the open work.
- Gradle emits an `archives` deprecation from the Kotlin Multiplatform plugin's
own `jvm()` target registration — upstream in Kotlin 2.2.10, not our build.
- Not built yet: LLM persona layer, opt-in coach, Compose UI.