6.3 KiB
6.3 KiB
Poker — Texas Hold'em with teachable AI opponents
Kotlin Multiplatform. Ships iOS + Android; Android first (only Android hardware for physical testing).
Build
No java/gradle on PATH — use Android Studio's bundled JDK:
export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
./gradlew :engine:jvmTest # engine tests (JVM)
./gradlew :engine:testAndroidHostTest # same suite, Android variant
./gradlew :sim:run --args="50000" # simulate 50k hands, print bot stats
./gradlew :sim:test # fast calibration-policy tests
./gradlew :sim:run --args="styles" # enforced controlled style experiment
./gradlew :sim:run --args="calibrate" # enforced 4×100k fixed-pool skill + style calibration
./gradlew :app:assembleDebug # build the APK
Layout
| Path | What |
|---|---|
engine/src/commonMain/.../core/ |
Cards, evaluator, equity, pre-flop chart |
engine/src/commonMain/.../bot/ |
Skill/style profiles, MathBot |
engine/src/commonMain/.../game/ |
Table — betting rounds, side pots, showdown |
sim/ |
JVM-only headless simulator used to tune bot profiles |
app/ |
Android app: Compose table, PokerViewModel |
assets/cards/ |
52 CC0 card faces + generated backs (source of truth) |
tools/generate_card_assets.sh |
Rasterises those SVGs into app/.../drawable-* |
engine is pure Kotlin with no platform APIs, so androidTarget() /
iosArm64() slot in without touching commonMain.
Design rules
- Poker maths never goes near the LLM. Difficulty and style are engine-side
EV/frequency calculations — instant, deterministic, testable, offline. The LLM
only narrates numbers the engine already computed (
DecisionTrace), and adds persona/table talk. - Skill and style are orthogonal.
SkillLevel= how correct decisions are;PlayStyle= bluffing, sandbagging, aggression, tightness. Build the strongest bot, then inject controlled error for lower tiers. - Pre-flop is range-based, not equity-based. All-in equity overvalues trash
(7-2o has ~35% vs one random hand but is unplayable).
PreflopChartranks the 169 starting hands soloosenessmeans "plays the top N%". - The simulator is how bots get tuned. Run it after any bot change, across
several seeds — a single seed will happily agree with a wrong conclusion.
The quick 50k run is diagnostic; only
calibrateis an enforced result. - A skill parameter must not smuggle in a style change. Several bugs came
from exactly this:
positionAwarenesssilently reduced hands played,potOddsRespectsystematically loosened weak players (which is a winning adjustment, so it inverted the gradient), and error direction overwrote style entirely. Skill should change how well a decision is made, not how loose or tight the player is. - Pot odds are the post-flop baseline. Never apply a blanket implied-odds discount: it is categorically wrong on the river, and future value on earlier streets must account for future costs and reverse implied odds before it is called an advantage.
- DecisionTrace is the coach contract. It records raw pot odds, the actual adjusted threshold, every adjustment, intended and chosen actions, and whether a skill error changed the decision. The coach explains these values; it does not reconstruct hidden bot logic.
Testing notes
Tabletakes aCardSource, soStackedDeck.of(holes, board)gives fully deterministic hands. Use it for any rule test.- The simulator gives the deck its own RNG, separate from each bot's. Never
share one: bots consume RNG proportional to their
equityIterations, so a shared stream means changing a profile silently changes the cards dealt. ./gradlew :sim:run --args="chart"dumps the starting-hand ranking.- Small samples lie. 1,000 hands is not enough to rank profiles. Use the paired 4×100k fixed-opponent calibration before accepting a skill change; it computes confidence bounds and exits nonzero when the contract fails.
Status
- Evaluator: verified exhaustively against published frequencies for all 2,598,960 five-card hands. ~24M evals/sec.
- Engine: chip-conserving. Covered by tests: side pots, uncalled-bet refunds, action order, malformed agent output, TDA Rule 47 (incomplete raises do not reopen betting, but several that cumulatively reach a full raise do), and TDA Rule 20 (odd chip to the first winner left of the button).
- Bots: controlled style calibration holds style constant against the same five opponents and deal seed. Rock is 10.3% VPIP, Maniac 67.5%; looseness ordering, Calling Station passivity (9.8% PFR, 0.27 AF), Maniac aggression, and PFR relationships all pass.
- Skill calibration pairs four 100k-hand seeds. Every candidate occupies the same seat against the same fixed opponent pool and deal seed. Advanced and Expert are allowed to overlap, but both must beat Intermediate and Intermediate must beat Beginner with a positive 95% lower confidence bound. Current lower bounds are +26.63, +16.19, and +24.71 bb/100 respectively.
- Expert differs by mechanism: it alone maintains opponent reads. The aggression prior is measured by the controlled neutral TAG experiment (0.229 observed, 0.22 configured), not selected because it looks plausible.
Rules invariants that are easy to get wrong
- Reopening betting cannot be a boolean.
Seat.lastActedAtBetrecords the bet level a player last acted at; betting reopens when `currentBet - lastActedAtBet= minRaiseSize`. Several short all-ins can reach that together.
PreflopChartpercentiles are weighted by combination counts (pair 6, suited 4, offsuit 12, total 1326), so "top 12%" means 12% of dealt hands, not 12% of the 169 classes.- Anything consuming
DecisionContext.historyacross hands must key offhandNumber. History is cleared each hand, so a size comparison silently drops events. - Known-imperfect: win-rate magnitudes are still ~10x realistic, and several profiles are looser than their labels (the Rock plays ~38% VPIP, should be ~12%). Tuning is the open work.
- Gradle emits an
archivesdeprecation from the Kotlin Multiplatform plugin's ownjvm()target registration — upstream in Kotlin 2.2.10, not our build. - Not built yet: LLM persona layer, opt-in coach, Compose UI.