Poker-Benchmark

byPokr AI

A frozen suite of 1,000 six-max, 100bb cash-game composite quizzes for evaluating model decisions against stored GTO answer data. Every listed model completed the same V1 quiz set.

78.5%
Top GTO score
0.066bb/decision
Lowest average EV loss
1,000
Quizzes in V1 benchmark
● 01 / leaderboard

Leaderboard

V1 · 1,000 Quizzes
DISPLAY METRIC:

All listed models completed the same 1,000 V1 quizzes. The ranking is sorted by the selected metric.

Model / FormatAverage GTO ScoreAvg GTO Score
0%25%50%75%100%

Evaluation Protocol: Quizzes are drawn deterministically from frozen composite manifests. The model follows its own action path across preflop, flop, turn, and river.

Ground-Truth Isolation: Model prompts strictly exclude action IDs, opponent hole cards, quality labels, EV loss values, or GTO score indicators.

Score Scope: Average GTO score and EV loss include scored decisions along each model’s action path.

● 02 / methodology

Benchmark Architecture

Evaluation flow 01 — 04
  1. 01

    Manifest

    Frozen set of 1,000 quizzes

    6 positions · 3 starting spots

  2. 02

    Answer-Hidden Prompt

    Legal choices match the run format

    No opponent cards or scores

  3. 03

    Dynamic Traversal

    Preflop → Flop → Turn → River

    Each model chooses its own path

  4. 04

    GTO Ground Truth

    Stored EV loss for each answer

    Quality-based GTO score

Frozen 1,000-Quiz Manifest

A deterministic seed fixes the quiz selection. The 1,000-quiz manifest balances six table positions (166–167 each) and three preflop starting spots: First Raise In (333), Facing Open (333), and Facing 3-Bet (334).

Trajectory-Driven Multi-Street Play

The model chooses its own path through each action tree. A preflop fold can end the hand; continuing can lead to flop, turn, and river decisions.

Quality- and EV-Based Scoring

Each chosen answer carries a quality label and EV loss. The reported GTO score combines its quality tier with normalized EV loss; average EV loss is reported separately in big blinds per scored decision.

● 03 / spots

Benchmark Spots Suite

Each card opens one decision sampled from the frozen V1 quiz set. These cards illustrate situations; they do not report per-spot model performance.

● 04 / comparative analysis

Run Insights · V1

Mean GTO Score by Street

Simple mean across all listed models. Later-street scores cover only decisions each model reached.

preflop75.8%
flop68.9%
turn65.4%
river62.7%
● 05 / citation

Citation Draft

This benchmark preview does not yet have a stable public URL. Update the citation when the benchmark page is published.

@misc{poker_benchmark_2026, title = {{Poker-Benchmark: Long-Horizon Game-Theoretic Poker Decisions for Frontier LLMs}}, author = {Pokr Research Team}, year = {2026}, note = {Unpublished benchmark preview; frozen 1000-quiz v1 manifest.} }