Skip to the report

Pokémon TCG AI Battle · strategy report

Policy × Deck:
Engineering a Competitive Pokémon TCG Agent

Belief-aware policy training, frozen-policy deck optimization, and controlled evaluation of the resulting competitive system.

One agent, two parts: a policy and its deck.

The policy chooses the moves; the exact 60-card deck defines the plan it can execute. This report follows how we built, trained, and tested that pair.

01 · video

How the agent works.

Evaluation · Approach & rationale · Technical soundness

How clearly is the chosen approach articulated, and how well is the rationale for the model and methods explained?

How original and technically sound is the proposed approach?

Contribution
We developed a Pokémon TCG agent that brings card mechanics, exact deck composition, and evolving beliefs about hidden opponent cards into a shared decision model. We trained the agent with self-play PPO, progressing from broad practice across decks to specialization on the submitted lists.
Inside the video · 7:33
Watch how the board, the agent’s own deck, and estimates of hidden cards feed into decisions and learning. This visual walkthrough explains how the components fit together and why we chose this design.

PTCG-RL policy architecture

≈200M parameters · viewer-specific belief · one stateless decision pass
  • Visible state
  • Own deck
  • Belief
  • Network
  • Output
01 · input

Viewer-safe decision state

Public gameboard · discards · prizes · rules

Own visible statehand · unseen cards · effects

Legal actionscurrent choice only

Acting player = “me” · other player = “opponent”
02 · compile

Pack the live tokens

18 in play32 own hand24 public hand64 discards12 prizes3 global / context159 legal options72 own unseen

card identity + static mechanics + live state

≤384 live rows + 4 scratch + PLAN + VALUE1,024-d · empty rows removed · no position embedding
03 · reason

Stateless Transformer ×17

1,024 width16 × 64 heads2,048 SwiGLU
LayerNorm → QKV → self-attention → residual addLayerNorm → SwiGLU → residual add

PLANVALUECTXOPTION

The full input is rebuilt for every decision; the network carries no hidden state.
04–05 · update context

Keep known deck and inferred deck separate

from TransformerPLAN + VALUE
Known exactlyOwn-deck reader

Exact submitted 60-card deck↓card features + copy counts↓4 heads × rank 32

emits own-deck delta
Inferred for this viewerOpponent-belief update

same-seat prior + public reveal floor↓count + hidden-zone logits↓legal copy domains + Σ = 60

emits current posterior + belief delta
own-deck delta Condition PLAN + VALUE once current-belief delta
Action and value gradients stop at the posterior probabilities.
06 · output

Policy head + value head

Policy headPLAN + CTXmasked OPTION pointer + STOPlegal action distributionMulti-action loop pick → refresh legal-action mask → pick again ↺ until STOP
Value headVALUE[lose, win] probabilitiesV(s) = P(win) − P(loss)

Seat 0 loopposterior₀(t)prior₀(t+1)↺ next seat 0 decision

Controller transport
CPU · no backpropagation through time
no cross-seat path

Seat 1 loopposterior₁(t)prior₁(t+1)↺ next seat 1 decision

actor distributionlegal pickengine actiongame transitionrebuild next viewer state ↺
Solid arrows move inference data. Purple arrows update and transport the viewer’s explicit belief between that seat’s decisions.

For each decision, the controller converts the viewer-safe state, legal actions, exact submitted 60-card deck, deterministic public-reveal floor and that seat’s previous opponent-card posterior into tokens. The stateless Transformer updates the opponent-count posterior, uses that updated belief for action and value, then projects it to exactly 60 expected copies and transports the detached result only to the same seat’s next decision. There is no RNN state; the seats never share belief memory;

Architecture + training recordObservation, belief, legal action selection, and the evidence behind both submitted specialists.
PTCG·RL Architecture

One board → one move

From a visible board
to one legal action.

At every decision, the controller rebuilds the acting player’s information set, updates that seat’s posterior over opponent deck composition, and scores only the legal actions supplied by the game engine.

Start with the 7:33 architecture walkthrough. It introduces the project’s main original contribution: a transformer conditioned on viewer-observable state whose explicit opponent-deck posterior is updated for each player, conditions the current decision, and can support bounded look-ahead with fallback without exposing hidden truth.
PTCG-RL policy architecture≈200M parameters · viewer-specific belief · one stateless decision pass
  • Visible state
  • Own deck
  • Belief
  • Network
  • Output
01 · input

Viewer-safe decision state

Public gameboard · discards · prizes · rules

Own visible statehand · unseen cards · effects

Legal actionscurrent choice only

Acting player = “me” · other player = “opponent”
02 · compile

Pack the live tokens

18 in play32 own hand24 public hand64 discards12 prizes3 global / context159 legal options72 own unseen

card identity + static mechanics + live state

≤384 live rows + 4 scratch + PLAN + VALUE1,024-d · empty rows removed · no position embedding
03 · reason

Stateless Transformer ×17

1,024 width16 × 64 heads2,048 SwiGLU
LayerNorm → QKV → self-attention → residual addLayerNorm → SwiGLU → residual add

PLANVALUECTXOPTION

The full input is rebuilt for every decision; the network carries no hidden state.
04–05 · update context

Keep known deck and inferred deck separate

from TransformerPLAN + VALUE
Known exactlyOwn-deck reader

Exact submitted 60-card deck↓card features + copy counts↓4 heads × rank 32

emits own-deck delta
Inferred for this viewerOpponent-belief update

same-seat prior + public reveal floor↓count + hidden-zone logits↓legal copy domains + Σ = 60

emits current posterior + belief delta
own-deck delta Condition PLAN + VALUE once current-belief delta
Action and value gradients stop at the posterior probabilities.
06 · output

Policy head + value head

Policy headPLAN + CTXmasked OPTION pointer + STOPlegal action distributionMulti-action loop pick → refresh legal-action mask → pick again ↺ until STOP
Value headVALUE[lose, win] probabilitiesV(s) = P(win) − P(loss)

Seat 0 loopposterior₀(t)prior₀(t+1)↺ next seat 0 decision

Controller transport
CPU · no backpropagation through time
no cross-seat path

Seat 1 loopposterior₁(t)prior₁(t+1)↺ next seat 1 decision

actor distributionlegal pickengine actiongame transitionrebuild next viewer state ↺
Solid arrows move inference data. Purple arrows update and transport the viewer’s explicit belief between that seat’s decisions.
01

Begin with the information set

It sees everything the player may know—and nothing the player may not.

The controller rebuilds a complete observation for every decision. That makes the network stateless: it does not need an RNN state or a recurrent hidden state from earlier turns.

Observed game state

Cards in play, hand, discard piles, public reveals, Prize slots without hidden identities, counts, turn flags, prompts, and history that is public now.

Legal actions

The engine supplies the actions actually available at this decision. The model never needs to invent the action set.

Exact own deck

The submitted 60-card list gives strategic context: what remains possible, which cards support the plan, and which candidate deck is being played.

Seat-specific opponent-deck posterior

Only this seat’s detached estimate of the opponent’s card counts returns from its preceding decision.

Public-reveal lower bounds

Publicly seen copies set hard lower bounds. Private hands, face-down Prize identities, hidden deck order, and exact opponent-deck truth never enter.

02

Turn the board into rows

A fixed-slot observation encoding makes a changing card game computable.

The encoder reserves 384 fixed slots. Occupied slots describe one observed fact or one legal action through aligned card IDs, attack IDs, 231 numerical features, and a live-slot mask. Populated slots are masked or packed before attention.

000–152visible play state
153–311legal options
312–383own unseen cards
0–48ownership · position · live Pokémon · card facts · zone · knowledge · Special Energy
49–143legal option type · conditions · selection type, relations, and bounds
144–201global state · persistent locks, prevention, costs, retaliation, delayed KO
202–230public action history · temporal effects that cross snapshot boundaries

Feature sheet. The model input uses exactly 231 viewer-safe float channels. Card ID, secondary card ID, attack ID, and mask remain separate aligned buffers. The range summary above is the complete inference-facing layout used here.

  1. Writeone fact into aligned fields
  2. Stackfacts at stable addresses
  3. Packonly live rows into tokens

Why this shape works: stable addresses make training predictable, while packing keeps attention focused on what is present. Card identity says what a row is; numerical features say what is happening to it now.

03

Run one forward pass per decision

A transformer shares one picture of the current decision.

All live tokens exchange context through self-attention. The same shared representation feeds the belief updater, actor, and critic; separate readers let each output ask a different question.

≈200M
parameter count
1,024
embedding width
17
transformer blocks
16
attention heads
2×
SwiGLU MLP ratio
6
learned query tokens · 4 scratch + plan + value

Why the diagram shows six learned query tokens. Four scratch tokens provide shared reasoning workspace. The plan token reads that workspace to form action context, while the value token reads it for the value head.

Permutation-tolerant attention

Useful because the number and order of live facts can change while their relationships matter.

No recurrent hidden state

Every action is auditable against one explicit input. Only the belief object is transported by the controller.

Shared trunk

Action, value, and belief learn a common language for the board instead of three disconnected representations.

04

Handle hidden cards honestly

Belief is a constrained estimate, not privileged information.

At the first decision, a frozen deck/metagame prior supplies plausible opponent card counts. After that, the model combines this seat’s previous posterior with the current public position and public-reveal lower bounds.

  1. Priorplausible card-count distribution
  2. Evidencecurrent viewer-safe state + public reveals
  3. Updatenew count posterior for possible opponent cards
  4. Projectdeterministically preserve exactly 60 expected copies
  5. Detachcarry only to this same seat’s next decision
Seat A beliefSeat A next turn
Seat B beliefSeat B next turn

The seats never share belief memory. True counts are used only in a training-only auxiliary loss on opponent card counts; they are never policy inputs at play time. The current posterior conditions policy and value once, then becomes a detached posterior carried to the same seat’s next decision.

06

Return action and value

One head chooses; the other judges.

Policy headWhich legal move—or legal sequence?

Scores only the offered action rows. Single-action prompts normally take the masked argmax at evaluation. Multi-action choices are decoded autoregressively: after each pick, the engine updates legality, the mask is refreshed, and STOP ends the sequence when allowed.

Value headWhat is the expected return?

Estimates outcome value from the same belief-conditioned representation. The game reward is simple: A win is +1, a loss −1, and a draw 0. It guides learning and evaluates search leaves; it is an estimate, not an observed win rate or a guarantee.

PerfectDou-inspired training signal. We use privileged opponent-deck truth during training to supervise the auxiliary belief-count loss. The policy and value heads still receive the same viewer-safe representation, and hidden deck truth is never an input during gameplay. This follows the same principle: perfect information can teach the model during training without being exposed to the policy at execution time.

07

Training record

One shared foundation became two submitted specialists.

Both policies inherited the same generalist checkpoint, then continued on their own submitted deck. The promotion chart preserves the recorded counters: the generalist trunk first, then each specialist on a zero-based training timeline.

Generalized modelSpecialized models7M82.52%18M72.07%165M68.95%450M67.68%519M62.33%689M62.17%Mega Lucario2.27B0M58.75%25M57.81%104M59.75%440M61.96%749M61.36%1.54B55.11%1.99B55.37%Dragapult851M0M65.00%32M68.75%132M65.70%419M58.67%
Shared foundation688.7M decisions

Broad-policy checkpoint inherited by both branches.

Primary · Lucario2.749B cumulative

Includes 1.662B decisions after the submitted exact 60 entered training.

Second · Dragapult1.493B cumulative

688.7M shared foundation + 804.458M specialist decisions.

Submission compression

196.7 MiB on disk. FP32 at inference.

The checkpoint was packaged to fit the submission limit. This is archive compression only; gameplay inference remains FP32.

Storage receipt

The model was compressed only for the submission archive. Gameplay inference still uses FP32.

Stored parameter precisionshare of stored weights
8-bit bulk · 79.8%12-bit input/action · 9.5%14-bit final blocks · 10.6%

Inference precision · FP32. The archive is restored for gameplay; the submitted agent does not run quantized inference.

The simplest accurate summary

See what is knowable. Estimate what is hidden. Search plausible futures. Choose only what is legal. Learn from the game that actually happened.

02 · training

Learn the game, then refine the policy.

Broad practice, then specialists for Lucario and Dragapult.

Evaluation · Approach & rationale · Technical soundness · Robustness

How clearly is the chosen approach articulated, and how well is the rationale for the model and methods explained?

How original and technically sound is the proposed approach?

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

We train PPO with opponent-deck belief across complete decks, then concentrate practice on the Lucario list. Broad practice teaches the rules and card interactions; specialization develops a plan for the exact 60-card list. Game outcomes train action selection and value estimation, while true opponent-card counts provide a training-only auxiliary loss for the separate belief estimate.

End-to-end learning process

From full card coverage to two submitted specialists.

The initial pool covered every card identity; PPO then refined one shared policy into deck-specific specialists before checkpoint selection and packaging.

  1. 01Cover the card universeBuild a generalist

    370 initial deck entries · all 1,267 card identities covered. Coverage-fill lists closed the gaps before specialization.

  2. 02Generate experiencePlay legal games

    The simulator returned legal trajectories, rewards, and terminal outcomes from both seats.

  3. 03Improve with PPOUpdate policy and value

    PPO used those outcomes to improve action selection and value estimation; clipped updates limited abrupt changes.

  4. 04Specialize the submitted listsTrain each exact 60

    Lucario and Dragapult continued from the shared policy while retaining field opponents.

  5. 05Select and packageCompare, freeze, package

    Seat-balanced evaluation selected the checkpoint before compression and submission.

Training deck dataset

370 deck entries · 356 unique lists

Base practice corpus; specialists retained varied opponents and added live competition lists. Counts show unique lists by source.

  • Synthetic legal coverage21 lists
  • Kaggle episode lists120 lists
  • Limitless tournaments107 lists
  • PokémonCard.io community108 lists
Why train across decks? Shared card and action representations let one policy learn across deck plans. Real decks supplied variety; synthetic lists filled card-coverage gaps. The reliability tests examine whether performance extends across opponents and conditions.
Expanded metagame gauntlet

16 tournament archetypes + one long-tail retention group

1,229 exact list variants, grouped by their archetype. The original 370-deck generalist corpus plus the additional lists from leading Kaggle competition decks. Each group receives 5.88% of learner-game starts; list count adds matchup variety within an archetype.

  • Dragapult ex134 lists
  • Crustle151 lists
  • Hydrapple ex27 lists
  • Slowking29 lists
  • Alakazam207 lists
  • Mega Kangaskhan ex15 lists
  • Mega Lucario ex80 lists
  • Rillaboom50 lists
  • Dudunsparce ex19 lists
  • Marnie's Grimmsnarl ex41 lists
  • N’s Zoroark ex16 lists
  • Hop's Trevenant58 lists
  • Archaludon ex47 lists
  • Team Rocket's Spidops44 lists
  • Ethan's Typhlosion35 lists
  • Teal Mask Ogerpon ex29 lists
  • Long-tail / out-of-meta retention247 lists

03 · concept and execution

Choose a deck the policy can execute.

Why Mega Lucario first, and Dragapult second?

Evaluation · Deck concept & strategy

How clearly is the deck concept articulated, and how well does it align with the intended strategy?

Why did Lucario lead, and why did we build Dragapult?

Lucario’s strong generalist policy–deck pairing justified a longer fixed-deck training run to study performance scaling with training decisions while joint policy–deck optimization ran in parallel. The joint loop did not produce a validated deck before the deadline. We then built Dragapult as a secondary deck chosen for setup consistency, with multiple outs to Basics, Evolution pieces, and Energy plus multiple setup and attack lines, reducing dependence on any one opening hand.

Primary submission

Lucario Hariyama

Lunatone–Solrock draw engine

Specialist practiceOn the submitted Lucario list
1.662Bdecisions
Final ladder ratingMean rating · 20–31 Aug
1,161.551,187.95
Captured game score700 wins · 2 draws · 520 losses
57.36%1,222 games
Training system
8 × RTX 5090 (VAST.ai)
Median update MFU
47.35%

Second submission

Dragapult Munkidori

Control and damage spread

Specialist practiceAcross the Dragapult specialist run
804.5Mdecisions
Final ladder ratingMean rating · 20–31 Aug
1,124.441,130.00
Captured game score703 wins · 1 draws · 519 losses
57.52%1,223 games
Training system
3 × RTX 5090 + 3 × RTX 4090 (local system)
Median update MFU
38.81%
Card counts → available options → observed play

Do the cards support the plan?

Opening-card probabilities explain the options the counts provide. Replay records show which draw, attack, and tactical options the policies used.

Primary submission

Lucario Hariyama

In a raw seven-card hand
Riolu or Pokémon search15 physical cards
88.25%
Mega Lucario or Ultra Ball8 physical cards
65.36%
Fighting Energy or Fighting Gong17 physical cards
91.66%
At least two Premium Power Pro4 physical cards
6.32%
Observed play · 290 replayed games
Lunar Cycle uses
1,004
Mega Lucario attacks
831
games using Premium Power Pro
237
Second submission

Dragapult Munkidori

In a raw seven-card hand
Dreepy or Pokémon search16 physical cards
90.08%
Drakloak or Pokémon search12 physical cards
80.94%
Dragapult ex or Ultra Ball7 physical cards
60.09%
Fire or Psychic Energy, or Crispin11 physical cards
77.76%
Observed play · 266 replayed games
Recon Directive uses
2,183
Dragapult ex attacks
817
games using Unfair Stamp
145

Each percentage describes a separate card-draw event. Shared search cards cannot fill every need at once. Card use shows that an option was exercised; it does not establish good timing or optimal counts.

Drawing the cards

From a 60-card deck, draw seven without replacement. All rows ask for at least one listed card, except Premium Power Pro, which asks for at least two of its four copies.

P(X ≥ 1) = 1 − C(60−k, 7) / C(60, 7)P(Premium ≥ 2) = 1 − [C(56, 7) + 4 C(56, 6)] / C(60, 7)

Here k is the listed physical-card count and C(n, r) counts combinations. These raw-hand baselines do not condition on a legal Basic Pokémon opening or model mulligans, Prizes, search discard costs, evolution timing, or shared search allocation.

Prizes are a separate calculation

In a random six-card Prize set, the chance that no stage of the main Evolution line is fully Prized is Lucario 99.94%; Dragapult 99.94%. This does not guarantee access to those cards during play.

Source and coverage

Card rules were checked against the frozen competition engine. The frozen card-use audit covers:

  • Lucario: 17/17 distinct card names used; 10,792 main decisions in 290 games.
  • Dragapult: 22/22 distinct card names used; 13,309 main decisions in 266 games.

Ability and attack totals count uses, which can repeat in one game. Premium Power Pro and Unfair Stamp totals count games with at least one use. These replay samples do not compare the decks under common match conditions.

Submitted policy–deck pair

Why Lucario was primary—and what Dragapult added.

Lucario’s strong generalist policy–deck pairing justified a longer fixed-deck training run to study performance scaling with training decisions. When the parallel joint optimizer did not yield a validated deck before the deadline, we added a secondary Dragapult deck chosen for setup consistency with multiple setup outs.

2.75B cumulative decisionsLucario specialist

1.66B after exact-list activation

→
156,800 evaluation gamesCheckpoint evaluation

selected among 10 Lucario checkpoints

→
160 / 160 actionsPackage fidelity

the packaged model reproduced all reference actions

→
Official #19 · mean #9Competition rank context

1161.55 final · 1183.89 mean rating

Full deck recordSubmitted deck plans, policy–deck crossover, optimizer trace, Codex proposals, and joint-optimization limits.

PTCG·RL Deck selection

Submitted pair · selection and research record

Why Lucario led
our submission.

The generalist policy had already scored Mega Lucario among its strongest deck pairings. We kept its training on the exact 60-card list running to study how performance scaled with training decisions while joint policy–deck optimization ran in parallel. When that search did not produce a validated deck before the deadline, we built Dragapult as a secondary deck chosen for setup consistency.

Competition decision

Mega Lucario / Hariyama was the primary submitted deck.

Lucario’s strong generalist policy–deck pairing justified a longer fixed-deck training run to study performance scaling with training decisions. The parallel joint optimizer did not yield a validated deck before the deadline, so we added a secondary Dragapult deck chosen for setup consistency with multiple setup outs.

#19 / #9
official finish / 23-team mean-rating order
1161.55 final / 1183.89 mean rating
57.76%
Lucario win rate over all captured games
826 wins / 1,430 games
55.18–60.30%
95% Wilson confidence interval; the public estimate can still vary with more games and a different opponent mix
156,800
checkpoint-selection games balanced by seat
160/160
packaged model reproduced the reference actions; verifies packaged-model fidelity
Specialist training traceFifty evaluations followed Lucario through seven policy promotions.
Mega Lucario gate scores across 50 evaluations, with seven policy promotions and a submitted-deck curriculum moving from 70 to 60 to 80 to 95 percent

Read: the lower strip is the share of learner games assigned to Lucario’s submitted deck; the remainder used field decks. Promotion resets make the blue line a within-generation progress trace, not one absolute score across the full run.

Historical checkpoint selection

This statistic selected within Lucario; it was not a portable win-rate baseline.

10 candidates · 156,800 games · 49 exact opposing decks · 10 opponent policies · both seats

81.34%
deck-equal game score
95% interval 78.24%–82.53%
79.93%
top-20-weighted game score
58.93%
worst archetype-family game score
68.16%
worst-deck-decile game score
4.25 pp
physical-seat gap
58.44%
posterior probability best among the 10 checkpoint candidates

The 81.34% figure is game score over a mixed ten-policy selection panel, where draws receive half credit. It is retained as a historical selection statistic, not reported as win rate and not used to claim that Lucario was stronger than Dragapult.

01

Submitted deck plans

Different decklists create different setup routes, resource priorities, and Prize trades.

Fighting Energy supports early pressure and key knockout thresholds. Dragapult was built for consistency: multiple outs to Basics, Evolution pieces, and Energy; Drakloak’s selective draw; and several setup and attack lines reduce dependence on any one opening hand.

2.75BLucario selected checkpoint · 1.66B on the submitted exact 60
1.49BDragapult cumulative training decisions · 688.7M generalist + 804.46M specialist
02

Policy × deck fit

The specialists were strongest on the decks they trained to play.

All three frozen policies played both submitted lists against the same frozen opposing policy and the same 35 fixed opponent decklists. Every cell contains 2,240 games with balanced seats.

Lucario deck
Dragapult deck
Generalist policy
48.62%transfer baseline
45.09%transfer baseline
Lucario policy
66.88%trained pairing
55.76%crossed pairing
Dragapult policy
56.34%crossed pairing
67.37%trained pairing
03

Frozen-policy deck optimizer

Archetype-guided optimizer

Deck search improved a bad list, but did not earn the submitted slot.

The frozen-policy optimizer began from a deliberately weak Dragapult list and drew from a candidate card pool built from known Dragapult-family lists, so this was archetype-guided rather than blind. Fresh evaluation showed a substantial recovery, yet the optimized list still lost to the policy’s own training list and missed the preregistered non-inferiority target.

Optimization trace

The full climb, with late progress magnified.

300 rounds · 1,041,408 search games
15,997 distinct decks evaluated

Full-range optimization traceThe trace begins at the deliberate 1.43 percent baseline, reaches a 20 percent best random screen at round zero, later moves around 60 percent, and peaks at 65.83 percent in round 265.FULL RANGE · 0–70%0%20%40%60%050100150200250300start · 1.43%best random screen · 20.0%best · 65.83%optimizer roundMagnified late optimization traceAn expanded linear scale from 56 to 67 percent shows the noisy accepted updates from rounds 200 to 300. The best later search estimate is 65.83 percent in round 265; a fresh held-out of that deck scored 58.20 percent.LATE SEARCH · EXPANDED 56–67%58%60%62%64%66%20025030064% earlier training-list testbest · 65.83%optimizer round
Initial deck1.43%
Optimized deck61.28%
Policy’s training list66.45%
04

Guided deck research tested · no replacement yet

Guided deck research

A proposal model prioritized candidates; fresh game results determined what advanced.

A persistent Codex goal kept the current matchup weakness, legal deck constraints, proposal history, and new results in one optimization loop. The language model used Pokémon TCG structure to prioritize plausible card-count changes, then revised its hypothesis after fresh fixed-policy games. The run produced useful proposals, but none of the five latest challengers met the deck-replacement criteria.

  1. GoalName the weakest matchup
  2. ProposeMake legal card-count changes
  3. EvaluatePlay fresh candidate/control games
  4. UpdateRetain outcomes; discard changes unsupported by fresh games
Proposal modelPrioritize legal candidate regions

Codex used matchup context, card roles and previous outcomes to rank which legal card-count changes deserved expensive simulation. The proposal score was not evidence of deck strength.

Acceptance authorityFresh games only

Codex used card knowledge and earlier game results to suggest which allowed deck changes to test. A deck was accepted only when fresh games supported it. This was guided search, not blind self-learning.

05

Joint optimization tested · no replacement yet

The loop strengthened a policy checkpoint, not a validated replacement deck.

The policy–deck training loop alternated legal card-count proposals with gameplay training on a fixed deck. Its best recorded policy reached a strong anchor result, but the independent terminal evaluation remained incomplete and the latest challengers did not meet the replacement criteria.

1.95B
environment steps at the best recorded policy promotion
82.19%
game score on its 800-game anchor; raw win rate 81.50%
41.25%
10th-percentile score across retained decks
0 / 5
candidate decks meeting the replacement criteria

Decision summary

Lucario was the scaling study. Dragapult was the consistency answer.

1 · SignalThe generalist policy already scored Lucario among its strongest deck pairings.
2 · ScaleWe extended the fixed-deck training run to study performance scaling with training decisions while joint optimization continued in parallel.
3 · AdaptWhen joint search did not yield a validated deck in time, Dragapult supplied multiple setup outs and multiple attack lines.

Deck optimization with a frozen policy

Evaluation · Key-card selection & use

How effectively are the key cards selected and utilized to support the deck's overall game plan?

Card counts must support a plan the policy can execute. We searched the Dragapult card pool while holding gameplay fixed, testing deck selection through outcomes.

Search method

Evolutionary search mutates and crosses over legal 60-card decks. A trainable surrogate predicts performance from game outcomes to prioritize simulations. A trained, frozen Dragapult policy plays both sides against 18 opponent decks. Final comparisons use fresh local games.

Candidate cards

Archetype-guided, not blind. 41 unique candidates: 40 from 56 known Dragapult-family lists, plus Grass Energy from the poor starting deck. Proposals used this run’s outcomes; the policy’s exact training list stayed hidden until selection was frozen.

Initial deck1.43%33 wins / 2,304 games
Optimized deck61.28%1,412 wins / 2,304 games
Policy’s training list66.45%1,531 wins / 2,304 games

Gain: +59.85 percentage points · paired schedule-block 90% resampling interval: +58.34 to +61.36. Win rate is wins / total games; draws are not wins. The policy’s training list was revealed only after selection. The round-265 search peak was 58.20% on a later matched held-out (1,341 / 2,304); paired schedule-block 90% resampling interval versus the selected deck: −5.51 to −1.61 pp.

The optimized deck still trails the policy’s training list by 5.16 points. The paired schedule-block 90% resampling interval for optimized minus training list is −7.22 to −3.11 points. Its one-sided 95% lower bound is −7.22 points, below the −5-point threshold, so the earlier campaign’s goal of showing a deficit smaller than 5 points was not established. These intervals pair 128 schedule blocks across 18 opponents; individual games are not paired.

The policy should receive its exact 60-card deck as input. This tells it which cards the current deck contains, so it need not rely on memorizing a training list to plan its play. This run used the legacy architecture without that input; whether adding it improves deck optimization remains untested.

Optimization trace

The full climb, with late progress magnified.

300 rounds · 1,041,408 search games
15,997 distinct decks evaluated

Full-range optimization traceThe trace begins at the deliberate 1.43 percent baseline, reaches a 20 percent best random screen at round zero, later moves around 60 percent, and peaks at 65.83 percent in round 265.FULL RANGE · 0–70%0%20%40%60%050100150200250300start · 1.43%best random screen · 20.0%best · 65.83%optimizer roundMagnified late optimization traceAn expanded linear scale from 56 to 67 percent shows the noisy accepted updates from rounds 200 to 300. The best later search estimate is 65.83 percent in round 265; a fresh held-out of that deck scored 58.20 percent.LATE SEARCH · EXPANDED 56–67%58%60%62%64%66%20025030064% earlier training-list testbest · 65.83%optimizer round
Card counts

Initial deck → optimized deck.

41 card-copy replacements
8 card-copy differences from the policy’s training list

All 41 unique candidate cards · three 60-card lists · zero counts included.Scroll horizontally to compare all three decks.

All 41 unique candidate cards in panels of 21 and 20 rows, with verified initial-deck, optimized-deck, and training-list counts, including zeros. Blue highlights the optimized deck.
The policy’s training list was revealed only after the optimizer selected its deck; the optimizer never received it. This boundary withheld the exact target counts; known archetype lists still supplied the candidate card pool.
More actions, less raw Energy

Energy fell from 41 to 13 cards, while Pokémon rose from 7 to 17 and Trainers from 12 to 30.

Dragapult became the core

The Dreepy–Drakloak–Dragapult ex line expanded from 1–1–1 to 4–4–3, and six Trainer cards reached four copies.

Every candidate is shown

Seventeen candidates have zero copies in all three lists. The other 24 account for every card used in the compared decks.

04 · performance

How did Lucario perform in the competition?

Evaluation · Competition performance

Performance within the competition track.

Official team finish: #19 of 6,807 teams.

Source: frozen Kaggle submission records, 18 August–1 September 2026. Each agent faced its own public match schedule; these are not head-to-head comparisons.

Submitted agentGamesWins / draws / lossesWin rateGame score
Lucario Hariyama1,222700 / 2 / 52057.28%57.36%
Dragapult Munkidori1,223703 / 1 / 51957.48%57.52%

Win rate = wins / games. Game score = (wins + ½ draws) / games.

Why do we propose Shared Rank, and how does it compare with other ranking methods?

We propose equal standing for teams in the same Shared Rank among the top 23.

On rank alone, they deserve equal awards; differences should be justified by other judging criteria.

We favor this view because it shows rating level, variation over time and alternative ranking estimates together. Shared ranks group overlapping middle-50% rating ranges.

Lucario SR3

12 teams share Shared Rank 3 (#SR3)

Mean rating
#9 / 231187.95 rating
Median rating
#10 / 231187.36 rating
Bayesian tournament
#11 / 10095% predictive: #6–#19
AR(1) forecast
#10 / 2395% scenario: #3–#20
Bradley–Terry
#9 / 10095% bootstrap: #8–#15
Official finish
#19 / 6,8071161.55 final rating

Mean and median: time-weighted ratings, 20–31 Aug 2026.

Scores across the top 23 teams

The top 23 teams are ordered by their average Kaggle rating over time, highest first. Adjacent teams have overlapping middle-50% rating ranges, grouped only while every member Q25–Q75 box has a common overlap. This proposed ranking gives teams in the same Shared Rank equal standing and equal claim to awards on rank alone; awarding more requires other judging criteria. The # beside each name is that team's mean-rating order among 23, numbered 1 through 23. The median, mean, official-final, and three model-reference markers share each row centerline. Potential Final Ranks Among the Top 23 Teams Top 23 teams · ordered by average Kaggle rating over time, highest first Shared rank, equal standing We propose equal standing for teams in the same Shared Rank among the top 23. On rank alone, they deserve equal awards; differences need other judging criteria. median mean official final Bayesian posterior-predictive tournament simulation rank · 100-team field Robust mean-reverting AR(1) rating-forecast rank · 23-team field Regularized batch Bradley–Terry paired-comparison rank · 100-team field MEAN # · TEAM · MEAN · σ · WIN RATE KAGGLE RATING · 1000–1400 1000 1000 1050 1050 1100 1100 1150 1150 1200 1200 1250 1250 1300 1300 1350 1350 1400 1400 SHARED RANK #SR1 1 team · Kaggle rating range 1313.18–1353.92 #1 Luca Mean 1335.0 · σ 25.2 Win rate 70.77% SHARED RANK #SR2 6 teams · Kaggle rating range 1203.65–1275.53 #2 palsystem Mean 1256.5 · σ 25.0 Win rate 59.69% #3 Unown Gradiant Mean 1255.3 · σ 28.5 Win rate 60.35% #4 flg Mean 1250.5 · σ 30.7 Win rate 59.99% #5 KawattaTaido Mean 1247.5 · σ 25.6 Win rate 58.93% #6 Petit Canard Mean 1231.1 · σ 20.8 Win rate 59.39% #7 LumenLiquidity Mean 1220.3 · σ 24.7 Win rate 58.67% SHARED RANK #SR3 12 teams · Kaggle rating range 1135.52–1218.38 #8 Azat Akhtyamov Mean 1194.5 · σ 27.4 Win rate 57.37% #9 @kdcyberdude Mean 1188.0 · σ 26.5 Win rate 57.76% #10 Majkel1337 Mean 1185.0 · σ 37.4 Win rate 59.83% #11 やる気元気ミワハルキ Mean 1181.4 · σ 23.2 Win rate 58.37% #12 LiamK Mean 1179.8 · σ 25.4 Win rate 57.43% #13 James Cox & Henry Chao Mean 1178.2 · σ 28.0 Win rate 58.67% #14 Sixth Sense Mean 1176.0 · σ 35.2 Win rate 58.18% #15 Klein Houmani Mean 1170.7 · σ 34.7 Win rate 58.07% #16 e-toppo + kurupical Mean 1167.8 · σ 47.0 Win rate 60.09% #17 Rmy Mean 1162.6 · σ 29.3 Win rate 57.04% #18 Preferred 213tubo Mean 1158.5 · σ 31.5 Win rate 57.66% #19 goonew Mean 1156.8 · σ 40.9 Win rate 56.90% SHARED RANK #SR4 3 teams · Kaggle rating range 1116.86–1156.14 #20 MissingNo. Mean 1145.0 · σ 17.6 Win rate 58.13% #21 Oshbocker Mean 1136.9 · σ 30.2 Win rate 57.43% #22 李秉叡(ntumlnoob) Mean 1131.4 · σ 25.6 Win rate 58.02% SHARED RANK #SR5 1 team · Kaggle rating range 1086.98–1124.21 #23 熱異常 Mean 1105.2 · σ 27.6 Win rate 58.48%

Download the shared-rank figure as SVG

How team ratings changed during competition

12 teams in Shared Rank #SR3 among the top 23. Choose “One team” to highlight a team alongside its group.

Preparing the rating path…
Top 23 teams: head-to-head game scores
07.2 · Field comparisons

How do the teams and their decks compare?

Evaluation · Competition performance · Robustness

Performance within the competition track.

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

Read across each row to find favorable and difficult matchups. The matrix shows game score against the other top-23 teams; the deck comparison shows similarity to published tournament lists.

Potential shared rank among 23 teams · one selected submission per team

Observed game score between every pair of selected teams

0%50%100%orange = below 50% · gray = even · green = above 50%printed rates are raw; small-sample colors fade toward even
DirectionRow team against column team
Percentage(wins + ½ draws) ÷ games
Cohort columnGames against the other 22 selected teams
Evidence5,601 Kaggle gameplay episodes
Result %among top 23
#SR11
#SR22
#SR23
#SR24
#SR25
#SR26
#SR27
#SR38
#SR39
#SR310
#SR311
#SR312
#SR313
#SR314
#SR315
#SR316
#SR317
#SR318
#SR319
#SR420
#SR421
#SR422
#SR523
TOP 23
ALL PUBLIC
#SR11Luca
—
90.4%n=104
67.3%n=110
61.7%n=107
52.0%n=75
46.1%n=52
66.7%n=45
71.4%n=21
54.5%n=11
63.6%n=11
60.0%n=5
75.0%n=8
37.5%n=8
66.7%n=3
100.0%n=2
100.0%n=10
80.0%n=5
75.0%n=4
66.7%n=3
0.0%n=1
—
100.0%n=1
—
66.7%n=586
70.6%n=991
#SR22palsystem
9.6%n=104
—
36.9%n=84
71.4%n=63
46.1%n=52
60.7%n=61
52.7%n=55
67.7%n=34
15.4%n=26
33.3%n=18
78.3%n=23
57.1%n=14
72.7%n=22
33.3%n=21
92.9%n=14
88.9%n=18
57.1%n=14
92.9%n=14
94.4%n=18
71.4%n=7
60.0%n=5
66.7%n=3
—
50.0%n=670
59.0%n=1278
#SR23Unown Gradiant
32.7%n=110
63.1%n=84
—
60.0%n=65
68.6%n=70
64.4%n=52
40.0%n=45
64.7%n=34
52.9%n=17
45.2%n=31
42.1%n=19
65.0%n=20
52.0%n=25
47.4%n=19
100.0%n=15
76.9%n=13
35.3%n=17
64.7%n=17
73.7%n=19
92.3%n=13
66.7%n=6
100.0%n=3
0.0%n=1
56.2%n=695
59.4%n=1312
#SR24flg
38.3%n=107
28.6%n=63
40.0%n=65
—
56.8%n=74
66.2%n=68
77.5%n=49
37.5%n=24
51.6%n=31
66.7%n=27
52.4%n=21
55.6%n=18
55.6%n=18
65.4%n=26
46.7%n=15
45.5%n=22
40.0%n=10
76.2%n=21
62.5%n=16
90.9%n=11
66.7%n=9
83.3%n=6
66.7%n=3
52.7%n=704
58.8%n=1268
#SR25KawattaTaido
48.0%n=75
53.9%n=52
31.4%n=70
43.2%n=74
—
63.8%n=47
43.6%n=55
55.6%n=27
76.0%n=25
21.4%n=14
67.9%n=28
70.8%n=24
36.8%n=19
82.7%n=26
77.8%n=18
72.2%n=18
66.7%n=6
47.1%n=17
43.5%n=23
75.0%n=4
36.4%n=11
0.0%n=4
0.0%n=1
51.6%n=638
58.6%n=1261
#SR26Petit Canard
53.9%n=52
39.3%n=61
35.6%n=52
33.8%n=68
36.2%n=47
—
53.1%n=49
61.3%n=31
93.1%n=29
29.2%n=24
85.2%n=27
82.6%n=23
33.3%n=24
89.5%n=19
75.0%n=24
47.4%n=19
82.3%n=17
43.5%n=23
33.3%n=15
66.7%n=12
36.4%n=11
43.3%n=15
100.0%n=3
51.8%n=645
58.9%n=1271
#SR27LumenLiquidity
33.3%n=45
47.3%n=55
60.0%n=45
22.4%n=49
56.4%n=55
46.9%n=49
—
47.1%n=34
46.1%n=26
68.2%n=22
55.6%n=27
71.4%n=21
65.5%n=29
47.8%n=23
13.0%n=23
21.1%n=19
45.8%n=24
56.3%n=16
68.2%n=22
81.8%n=11
70.0%n=10
62.5%n=16
100.0%n=2
49.1%n=623
56.9%n=1243
#SR38Azat Akhtyamov
28.6%n=21
32.4%n=34
35.3%n=34
62.5%n=24
44.4%n=27
38.7%n=31
52.9%n=34
—
26.9%n=26
80.0%n=35
59.4%n=32
73.1%n=26
47.6%n=21
41.9%n=31
15.0%n=20
28.6%n=21
52.4%n=21
63.6%n=22
53.6%n=28
38.1%n=21
50.0%n=16
78.6%n=14
50.0%n=6
47.9%n=545
56.5%n=1238
#SR39@kdcyberdude
45.5%n=11
84.6%n=26
47.1%n=17
48.4%n=31
24.0%n=25
6.9%n=29
53.9%n=26
73.1%n=26
—
79.3%n=29
40.7%n=27
45.8%n=24
51.6%n=31
52.0%n=25
75.0%n=24
34.8%n=23
38.9%n=18
56.3%n=16
73.1%n=26
29.2%n=24
79.0%n=19
30.8%n=13
57.1%n=7
51.5%n=497
57.4%n=1222
#SR310Majkel1337
36.4%n=11
66.7%n=18
54.8%n=31
33.3%n=27
78.6%n=14
70.8%n=24
31.8%n=22
20.0%n=35
20.7%n=29
—
18.5%n=27
21.1%n=19
60.9%n=23
55.0%n=20
36.0%n=25
9.1%n=22
68.8%n=16
37.5%n=16
57.1%n=14
87.5%n=16
85.7%n=14
100.0%n=14
25.0%n=4
45.6%n=441
59.1%n=1272
#SR311やる気元気ミワハルキ
40.0%n=5
21.7%n=23
57.9%n=19
47.6%n=21
32.1%n=28
14.8%n=27
44.4%n=27
40.6%n=32
59.3%n=27
81.5%n=27
—
51.9%n=27
29.6%n=27
69.6%n=23
19.1%n=21
25.0%n=20
56.0%n=25
48.4%n=31
45.5%n=33
47.4%n=19
73.7%n=19
59.3%n=27
53.9%n=13
46.3%n=521
57.4%n=1258
#SR312LiamK
25.0%n=8
42.9%n=14
35.0%n=20
44.4%n=18
29.2%n=24
17.4%n=23
28.6%n=21
26.9%n=26
54.2%n=24
79.0%n=19
48.1%n=27
—
55.9%n=34
43.3%n=30
67.9%n=28
65.4%n=26
50.0%n=24
41.7%n=24
41.4%n=29
72.2%n=18
38.9%n=18
46.1%n=13
61.5%n=13
46.6%n=481
57.2%n=1212
#SR313James Cox & Henry Chao
62.5%n=8
27.3%n=22
48.0%n=25
44.4%n=18
63.2%n=19
66.7%n=24
34.5%n=29
52.4%n=21
48.4%n=31
39.1%n=23
70.4%n=27
44.1%n=34
—
45.5%n=33
7.7%n=26
80.0%n=30
38.9%n=18
44.4%n=27
60.0%n=25
64.7%n=17
37.5%n=8
66.7%n=9
64.3%n=14
49.6%n=488
58.0%n=1254
#SR314Sixth Sense
33.3%n=3
66.7%n=21
52.6%n=19
34.6%n=26
17.3%n=26
10.5%n=19
52.2%n=23
58.1%n=31
48.0%n=25
45.0%n=20
30.4%n=23
56.7%n=30
54.5%n=33
—
96.7%n=30
83.3%n=18
55.2%n=29
75.0%n=16
70.8%n=24
43.8%n=16
81.3%n=16
25.0%n=12
100.0%n=5
53.9%n=465
58.0%n=1280
#SR315Klein Houmani
0.0%n=2
7.1%n=14
0.0%n=15
53.3%n=15
22.2%n=18
25.0%n=24
87.0%n=23
85.0%n=20
25.0%n=24
64.0%n=25
81.0%n=21
32.1%n=28
92.3%n=26
3.3%n=30
—
67.7%n=34
84.2%n=19
95.7%n=23
35.0%n=20
27.8%n=18
100.0%n=19
59.1%n=22
85.7%n=14
54.2%n=454
57.8%n=1269
#SR316e-toppo + kurupical
0.0%n=10
11.1%n=18
23.1%n=13
54.5%n=22
27.8%n=18
52.6%n=19
79.0%n=19
71.4%n=21
65.2%n=23
90.9%n=22
75.0%n=20
34.6%n=26
20.0%n=30
16.7%n=18
32.4%n=34
—
17.6%n=17
94.1%n=17
18.2%n=11
68.8%n=16
25.0%n=12
55.6%n=18
33.3%n=12
45.7%n=416
59.5%n=1226
#SR317Rmy
20.0%n=5
42.9%n=14
64.7%n=17
60.0%n=10
33.3%n=6
17.6%n=17
54.2%n=24
47.6%n=21
61.1%n=18
31.3%n=16
44.0%n=25
50.0%n=24
61.1%n=18
44.8%n=29
15.8%n=19
82.3%n=17
—
57.1%n=35
55.0%n=20
50.0%n=20
57.1%n=14
42.0%n=25
62.5%n=16
49.1%n=410
57.5%n=1235
#SR318Preferred 213tubo
25.0%n=4
7.1%n=14
35.3%n=17
23.8%n=21
52.9%n=17
56.5%n=23
43.8%n=16
36.4%n=22
43.8%n=16
62.5%n=16
51.6%n=31
58.3%n=24
55.6%n=27
25.0%n=16
4.3%n=23
5.9%n=17
42.9%n=35
—
72.7%n=11
75.0%n=20
41.2%n=17
91.7%n=24
60.0%n=5
45.2%n=416
57.3%n=1247
#SR319goonew
33.3%n=3
5.6%n=18
26.3%n=19
37.5%n=16
56.5%n=23
66.7%n=15
31.8%n=22
46.4%n=28
26.9%n=26
42.9%n=14
54.5%n=33
58.6%n=29
40.0%n=25
29.2%n=24
65.0%n=20
81.8%n=11
45.0%n=20
27.3%n=11
—
76.9%n=13
62.5%n=16
75.0%n=12
50.0%n=8
46.3%n=406
56.7%n=1239
#SR420MissingNo.
100.0%n=1
28.6%n=7
7.7%n=13
9.1%n=11
25.0%n=4
33.3%n=12
18.2%n=11
61.9%n=21
70.8%n=24
12.5%n=16
52.6%n=19
27.8%n=18
35.3%n=17
56.3%n=16
72.2%n=18
31.3%n=16
50.0%n=20
25.0%n=20
23.1%n=13
—
30.0%n=20
58.6%n=29
60.0%n=10
41.4%n=336
57.0%n=1227
#SR421Oshbocker
—
40.0%n=5
33.3%n=6
33.3%n=9
63.6%n=11
63.6%n=11
30.0%n=10
50.0%n=16
21.1%n=19
14.3%n=14
26.3%n=19
61.1%n=18
62.5%n=8
18.8%n=16
0.0%n=19
75.0%n=12
42.9%n=14
58.8%n=17
37.5%n=16
70.0%n=20
—
75.0%n=12
46.1%n=13
42.8%n=285
57.7%n=1234
#SR422李秉叡(ntumlnoob)
0.0%n=1
33.3%n=3
0.0%n=3
16.7%n=6
100.0%n=4
56.7%n=15
37.5%n=16
21.4%n=14
69.2%n=13
0.0%n=14
40.7%n=27
53.9%n=13
33.3%n=9
75.0%n=12
40.9%n=22
44.4%n=18
58.0%n=25
8.3%n=24
25.0%n=12
41.4%n=29
25.0%n=12
—
57.9%n=19
40.2%n=311
58.2%n=1233
#SR523熱異常
—
—
100.0%n=1
33.3%n=3
100.0%n=1
0.0%n=3
0.0%n=2
50.0%n=6
42.9%n=7
75.0%n=4
46.1%n=13
38.5%n=13
35.7%n=14
0.0%n=5
14.3%n=14
66.7%n=12
37.5%n=16
40.0%n=5
50.0%n=8
40.0%n=10
53.9%n=13
42.1%n=19
—
40.8%n=169
58.0%n=1216
Deck-list context · Limitless Standard results

22 of 23 final-cohort lists were within eight card replacements.

Distance counts the fewest card-copy replacements needed to turn a selected competition list into its nearest exact-mappable list in the published Limitless Standard comparison pool.

23/23cohort lists resolved
3median replacements
16/23within four replacements
22/23within eight replacements
Nearest published list for each teamLimitless Standard

Fewest card-copy replacements from each selected 60 to its closest exact-mappable list in the published Limitless Standard pool before September 2026.

submitted to publishedΔ card-copy replacements
#SR11 team
#SR26 teams
#SR312 teams
#SR43 teams
#SR51 team

04 · consistency

How consistent were the results?

First repeat a fixed match schedule, then compare results across the competition.

Evaluation · Consistency

How consistently does the model perform under repeated matches and stable conditions?

05.1 · Consistency

Repeat the same match schedule.

Evaluation · Consistency

How consistently does the model perform under repeated matches and stable conditions?

Start by changing the shuffle while keeping the comparison panel fixed. A fixed-deck checkpoint is a saved policy evaluated while its assigned exact 60-card list remains unchanged.

Both agents’ results varied by a similar amount.

Dragapult won more often against this lineup. Its results fluctuated about as much as Lucario’s when we played fresh games.

What did we repeat?

  1. Keep the players fixed

    Same trained agent, own deck, opponent decks and opposing agents. No learning between games.

  2. Play 184 fresh games

    Use the same opponent schedule and equal first/second play, with new shuffles and chance events.

  3. Compare 100 batches

    Count wins in each batch of 184 games. That gives 100 win rates per agent, from 18,400 games each.

How far did the results move?

One dot = one batch of 184 games. Dots further right mean more wins. A wider spread means more variation between batches.

Average win rate 50% = 92 of 184 games won
Lucario–Hariyama
Average 55.5%

Observed range: 47.8–63.0% · 88–116 wins out of 184

Dragapult
Average 69.1%

Observed range: 61.4–77.2% · 113–142 wins out of 184

Win rate in one 184-game batch →
Win rate = wins ÷ all games; draws add no wins. Both rows use the same scale. Dots stack only to stay visible; their height has no additional meaning.

Consistency: similar variation

Both sets of results have a similar spread. A single 184-game batch could give a noticeably higher or lower win rate even though the agent had not changed.

Performance: a higher average

Dragapult’s results sit further right. It stayed above 50% in every batch because its results were centered higher, with a similar amount of fluctuation.

This measures consistency against our fixed local opponent lineup. It does not establish the same results against every competition opponent or across independently trained models.

Did Lucario’s results change during competition?

Evaluation · Consistency

How consistently does the model perform under repeated matches and stable conditions?

We found no clear change in Lucario’s results.

Against the same opponents in recorded Kaggle games, the balanced win rate was 52.56% early and 52.80% late. The data still allow a rise or fall, so this does not prove performance stayed unchanged.

Compare the two halves of the competition

We compare the same Lucario agent against the same 44 opponent agents from 27 teams. The balanced comparison gives each team equal weight and counts first and second turns equally.

Win rate · – UTC · two equal halves
ComparisonFirst half20 Aug–26 AugSecond half26 Aug–1 Sept
All Kaggle gamesOpponents and turn order vary57.69%330/572 wins57.12%329/576 wins
Same opponent agentsOnly agents faced in both halves50.12%216/431 wins52.76%239/453 wins
Same opponents, balanced comparisonEqual team weights; first and second turns count equally52.56%52.80%

Estimated change in the balanced win rate: +0.23 percentage points. Its 95% range is −10.63 to +10.50 points.

Two other comparisons also found no clear change

These checks group the same Kaggle records differently. Each change compares the last period’s balanced win rate with the first. Both ranges include zero, allowing either a rise or fall.

Same teams, allowing different agents

952 games · 82.9% of this window

−1.43 percentage points

95% range: −11.03 to +7.33 points

Same agents, four shorter periods

435 games · 37.9% of this window

+3.94 percentage points

95% range: −11.03 to +17.95 points

05 · robustness

How did results vary across opponents and starting conditions?

We examined Lucario’s Kaggle results for difficult matchups and dependence on favorable starts.

Evaluation · Robustness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

05.2 · Transfer

Does the advantage survive deck edits?

Evaluation · Robustness · Technical soundness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

How original and technically sound is the proposed approach?

Test whether the specialists’ learned play still helps when the exact card counts change.

Kaggle game scores by archetype · top 23 teams

Lucario’s matchup spread by opposing archetype

Each bar pools the selected Lucario submission's public games by the opponent's submitted archetype; n is games. This is descriptive matchup evidence, not a causal diagnosis.

84.6%
79.3%
75.0%
73.1%
54.6%
54.6%
53.9%
51.6%
40.7%
34.8%
20.9%
The clearest weakness was Slowking / Kangaskhan: Lucario scored 20.88% over 91 games, and every one of the four exact Slowking policies beat it individually. Results combine deck and policy; archived action replays cover only about 45% of gold-cohort games, so the proposed mechanism remains a testable explanation rather than a causal conclusion.
Kaggle games · four Lucario matchup loops

Can a strong overall score hide a weak matchup?

A → B means A had the higher game score. Reverse the direction for the weakness: B’s game score is 100% minus A’s. Every arrow in these four loops is above 70% game score.

Game score = (wins + ½ draws) / games n = games50% = even matchup

Loop 1, 3 archetypes. Lucario–Hariyama is highlighted in pale blue. Alakazam / Dudunsparce → Slowking / Kangaskhan: 82.3% game score, n=68; Slowking / Kangaskhan → Lucario–Hariyama: 79.1% game score, n=91; Lucario–Hariyama → Alakazam / Dudunsparce: 79.3% game score, n=29.Loop 2, 5 archetypes. Lucario–Hariyama is highlighted in pale blue. Alakazam / Dudunsparce → Slowking / Kangaskhan: 82.3% game score, n=68; Slowking / Kangaskhan → Lucario–Hariyama: 79.1% game score, n=91; Lucario–Hariyama → Crustle / Kangaskhan: 84.6% game score, n=26; Crustle / Kangaskhan → Dragapult / Moltres: 78.3% game score, n=23; Dragapult / Moltres → Alakazam / Dudunsparce: 81.5% game score, n=27.Loop 3, 5 archetypes. Lucario–Hariyama is highlighted in pale blue. Alakazam / Dudunsparce → Slowking / Kangaskhan: 82.3% game score, n=68; Slowking / Kangaskhan → Lucario–Hariyama: 79.1% game score, n=91; Lucario–Hariyama → Espeon / Kangaskhan: 75.0% game score, n=24; Espeon / Kangaskhan → Dragapult / Moltres: 81.0% game score, n=21; Dragapult / Moltres → Alakazam / Dudunsparce: 81.5% game score, n=27.Loop 4, 6 archetypes. Lucario–Hariyama is highlighted in pale blue. Alakazam / Dudunsparce → Slowking / Kangaskhan: 82.3% game score, n=68; Slowking / Kangaskhan → Lucario–Hariyama: 79.1% game score, n=91; Lucario–Hariyama → Espeon / Kangaskhan: 75.0% game score, n=24; Espeon / Kangaskhan → Raging Bolt / Ogerpon: 92.3% game score, n=26; Raging Bolt / Ogerpon → Froslass / Bronzong: 80.0% game score, n=30; Froslass / Bronzong → Alakazam / Dudunsparce: 90.9% game score, n=22.

Slowking was a concentrated weakness. In the selected top-23 games, Lucario’s game score was 58.4% against the other ten archetypes (406 games), but 20.9% against Slowking (91 games). That gap makes Slowking a clear priority for matchup-specific testing.

Which archetypes had the widest matchup coverage?

Count the opposing archetypes against which each had more than 50% game score. Each was compared with all eleven other archetypes in the selected top-23 Kaggle record.

  1. Lucario–Hariyama8 / 11
  2. Crustle / Kangaskhan7 / 11
  3. Dragapult / Munkidori7 / 11
  4. Hydrapple / Ogerpon7 / 11
  5. Dragapult / Blaziken6 / 11
  6. Espeon / Kangaskhan6 / 11
  7. Froslass / Bronzong6 / 11
  8. Alakazam / Dudunsparce5 / 11
  9. Arboliva ex / Ogerpon4 / 11
  10. Raging Bolt / Ogerpon4 / 11
  11. Dragapult / Moltres3 / 11
  12. Slowking / Kangaskhan3 / 11

Lucario had the broadest coverage: 8 favorable matchups; the next group had 7. Slowking and Dragapult / Moltres had the fewest, at 3 each. Yet Slowking was Lucario’s hardest opponent: a narrow counter can still expose a major weakness.

These counts compare the observed team-and-deck combinations. Matchup sample sizes and player skill vary; the matrix below shows the individual game counts and uncertainty.

Matchup matrixReference
Every final-cohort archetype against every other observed archetype

Which matchups were actually favorable?

The number in each cell is the archetype family's game score: a win counts as 1, a draw as ½. Color uses the neutral-shrunk estimate, so a tiny sample does not look more certain than it is.

Archetype family010203040506070809101112
01Espeon / Kangaskhan—64%n=8825%n=2449%n=977%n=1492%n=2687%n=2334%n=8235%n=2081%n=2168%n=3464%n=25
02Dragapult / Munkidori36%n=88—45%n=10849%n=53457%n=22953%n=10860%n=16555%n=54354%n=8753%n=11646%n=9672%n=108
03Lucario–Hariyama75%n=2455%n=108—55%n=8685%n=2652%n=3154%n=2621%n=9173%n=2641%n=2735%n=2379%n=29
04Hydrapple / Ogerpon51%n=9751%n=53445%n=86—60%n=12453%n=9844%n=10452%n=36660%n=8737%n=9978%n=7239%n=85
05Crustle / Kangaskhan93%n=1443%n=22915%n=2640%n=124—73%n=2253%n=5555%n=12394%n=1878%n=2389%n=1833%n=18
06Raging Bolt / Ogerpon8%n=2647%n=10848%n=3147%n=9827%n=22—34%n=2965%n=6960%n=2570%n=2780%n=3039%n=23
07Dragapult / Blaziken13%n=2340%n=16546%n=2656%n=10447%n=5566%n=29—56%n=13168%n=2256%n=2721%n=1968%n=22
08Slowking / Kangaskhan66%n=8245%n=54379%n=9149%n=36645%n=12335%n=6944%n=131—33%n=6362%n=10149%n=7118%n=68
09Arboliva ex / Ogerpon65%n=2046%n=8727%n=2640%n=876%n=1840%n=2532%n=2267%n=63—55%n=3382%n=1143%n=14
10Dragapult / Moltres19%n=2147%n=11659%n=2763%n=9922%n=2330%n=2744%n=2738%n=10145%n=33—25%n=2081%n=27
11Froslass / Bronzong32%n=3454%n=9665%n=2322%n=7211%n=1820%n=3079%n=1951%n=7118%n=1175%n=20—91%n=22
12Alakazam / Dudunsparce36%n=2528%n=10821%n=2961%n=8567%n=1861%n=2332%n=2282%n=6857%n=1419%n=279%n=22—
Observed result orderingReference
Observed result ordering

Opponent schedules differ by family, so this orders the archived cross-family results. It is not a schedule-balanced deck-strength ranking or a deck-only tier list.

Rank · familyGame scoreStable edgesBest / worst observed matchup
#1Espeon / Kangaskhann=45454.2%4 favorable2 unfavorableRaging Bolt / Ogerpon · 92.3%Crustle / Kangaskhan · 7.1%
#2Dragapult / Munkidorin=218253.2%2 favorable1 unfavorableAlakazam / Dudunsparce · 72.2%Espeon / Kangaskhan · 36.4%
#3Lucario–Hariyaman=49751.5%4 favorable1 unfavorableCrustle / Kangaskhan · 84.6%Slowking / Kangaskhan · 20.9%
#4Hydrapple / Ogerponn=175251.2%2 favorable1 unfavorableFroslass / Bronzong · 77.8%Dragapult / Moltres · 37.4%
#5Crustle / Kangaskhann=67050.0%1 favorable2 unfavorableArboliva ex / Ogerpon · 94.4%Lucario–Hariyama · 15.4%
#6Raging Bolt / Ogerponn=48849.6%2 favorable1 unfavorableFroslass / Bronzong · 80.0%Espeon / Kangaskhan · 7.7%
#7Dragapult / Blazikenn=62349.1%0 favorable2 unfavorableArboliva ex / Ogerpon · 68.2%Espeon / Kangaskhan · 13.0%
#8Slowking / Kangaskhann=170847.8%3 favorable3 unfavorableLucario–Hariyama · 79.1%Alakazam / Dudunsparce · 17.7%
#9Arboliva ex / Ogerponn=40646.3%1 favorable1 unfavorableFroslass / Bronzong · 81.8%Crustle / Kangaskhan · 5.6%
#10Dragapult / Moltresn=52146.3%2 favorable3 unfavorableAlakazam / Dudunsparce · 81.5%Espeon / Kangaskhan · 19.0%
#11Froslass / Bronzongn=41645.7%1 favorable2 unfavorableAlakazam / Dudunsparce · 90.9%Crustle / Kangaskhan · 11.1%
#12Alakazam / Dudunsparcen=44145.6%1 favorable4 unfavorableSlowking / Kangaskhan · 82.3%Froslass / Bronzong · 9.1%
A specific limit to Lucario’s robustness

How often did Lucario beat Slowking?

Evaluation · Robustness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

Kaggle · public match results

Lucario won about one in four games against Slowking.

Across the same public archive, Lucario won 25.58% of games against Slowking / Kangaskhan lists, compared with 57.28% overall. In the official final top 23, 4 teams used Slowking / Kangaskhan in their leading submission, including 2 in the top 10.

Kaggle · 18 Aug – 1 Sep25.58%
Lucario win rate against Slowking

66 wins, 1 draw and 191 losses in 258 games against Slowking / Kangaskhan lists.

These games span 14 opposing teams and 19 submissions.

Same Kaggle archive · all opponents57.28%
Lucario overall win rate

1,222 games against 188 other teams. The 258 Slowking games are included in this total.

Both percentages count wins / games; draws receive zero win credit.

The public results combine the effects of the decks, the opposing agents and the game situations. They do not isolate a particular Lucario decision error or show which training change would improve this matchup.

05.4 · Situational advantages

Examine performance after an early lead.

Follow the game from turn order and the opening hand to an early Prize lead.

How did turn order and starting hands affect results?

Evaluation · Robustness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

Both submitted agents won more often going first on Kaggle.

Lucario won 62.10% going first versus 53.32% going second. Dragapult won 61.09% versus 53.72%. Both differences remained positive when comparing games against the same opposing agents.

Kaggle public games · 16–31 August 2026. We checked all 2,832 recorded games for the two submitted agents. Each used its own fixed 60-card deck against independently developed opponents.

Lucario
1,430 Kaggle games

+8.78 percentage points going first

Win rate by actual turn order
  1. Played first62.10%
    449/723 wins
  2. Played second53.32%
    377/707 wins

First minus second · 95% range +3.67 to +13.88 points

+7.77 points against the same opposing agents95% range +2.44 to +13.10 points · 1,206 games against 71 agents faced in both turn orders

Dragapult
1,402 Kaggle games

+7.37 percentage points going first

Win rate by actual turn order
  1. Played first61.09%
    438/717 wins
  2. Played second53.72%
    368/685 wins

First minus second · 95% range +2.20 to +12.53 points

+7.17 points against the same opposing agents95% range +1.47 to +12.87 points · 1,139 games against 123 agents faced in both turn orders

The first-player advantage was present in both public records. Turn order was chosen during setup, so these are observed differences; the same-opponent comparison reduces differences in opponent mix.

Which opening cards were linked to wins on Kaggle?

Evaluation · Robustness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

Lucario won more often when its opening hand contained Fighting Energy or Fighting Gong.

Its win rate was 58.84% with at least one, versus 47.41% with neither — a 11.43-point difference. The difference was still positive after accounting for the opposing agent and turn order.

These are the seven cards visible to the agent before it chose its Active Pokémon. Each row compares games with the named cards against games without that condition. We show four comparisons for each deck. Dragapult’s rows cover Budew alone, Bench setup, evolution support and Energy disruption.

Reading the results: positive means a higher win rate with the condition; negative means lower. A 95% range crossing zero leaves either direction possible. The last column compares games against the same opposing agent and in the same turn order.

Lucario Hariyama1,430 Kaggle games · 222 opposing teams

Energy or Fighting Gong showed the clearest positive difference among the tested hand conditions. With the opposing agent and turn order held fixed, the estimated difference was +10.96 points (95% range +2.08 to +19.84). The Riolu + Mega Lucario + Fighting Energy group also had a higher observed win rate, but its range includes zero.

Opening-hand conditionWin rate with conditionWin rate without conditionDifferenceIndividual 95% range · pointsSame opponent & turn orderDifference · 95% range
Fighting Energy58.72%687/1170 wins53.46%139/260 wins+5.26−1.43 to +11.95+6.46−0.41 to +13.321,042 matching games
Fighting Energy or Fighting Gong58.84%762/1295 wins47.41%64/135 wins+11.43+2.59 to +20.28+10.96+2.08 to +19.84753 matching games
Mega Lucario or Ultra Ball58.72%515/877 wins56.24%311/553 wins+2.48−2.78 to +7.75−1.12−6.71 to +4.471,192 matching games
Riolu + Mega Lucario + Fighting Energy64.07%107/167 wins56.93%719/1263 wins+7.14−0.63 to +14.92+6.26−1.73 to +14.24921 matching games
Dragapult Munkidori1,402 Kaggle games · 281 opposing teams

Budew with Bench and evolution support showed a larger, uncertain difference.

This combination won 67.59% versus 56.33% without the full combination: the highest observed win rate among the setup combinations tested. After accounting for opponent and turn order, the gap was +14.79 points. Its corrected 95% range was −0.36 to +29.94 points, so a positive advantage remains uncertain.

Why test this plan? Budew’s zero-Energy Itchy Pollen attack blocks the opponent’s Item cards for their next turn. This can buy time to build the Bench, evolve Dreepy into Drakloak and draw with Recon Directive before attacking with Dragapult. Itchy Pollen does not directly prevent attacks. This follows Pokémon’s Budew strategy discussion and the card rules in the competition engine.

Combination definitions: a Bench starter is Dreepy or Buddy-Buddy Poffin; evolution support is Drakloak, Poké Pad or Dawn. “+” requires every group in the opening seven. Search cards indicate options, not guaranteed execution; Item lock, Prizes and timing can still prevent the sequence.

Opening-hand conditionWin rate with conditionWin rate without conditionDifferenceIndividual 95% range · pointsSame opponent & turn orderDifference · corrected 95% range
Budew60.00%243/405 wins56.47%563/997 wins+3.53−2.15 to +9.21+5.02−4.93 to +14.97953 matching games
Budew + Buddy-Buddy Poffin63.87%99/155 wins56.70%707/1247 wins+7.17−0.88 to +15.22+9.99−4.37 to +24.35598 matching games
Budew + Bench starter + evolution support67.59%98/145 wins56.33%708/1257 wins+11.26+3.16 to +19.36+14.79−0.36 to +29.94604 matching games
Budew + Crushing Hammer56.46%83/147 wins57.61%723/1255 wins−1.15−9.62 to +7.33−4.24−18.77 to +10.29562 matching games

Did the agent use the plan? It used Itchy Pollen by its second own turn in 136/145 games with the Budew + Bench + evolution combination; all 136 had a Dreepy-line Pokémon on the Bench when it used the attack. Starting without Budew did not mean playing without it: the agent still used early Itchy Pollen in 649/997 such games (65.10%). These are recorded choices, not evidence that the attack caused the wins.

Situational advantage · public replay evidence

Taking at least two Prizes on a player’s second turn was strongly associated with winning.

Evaluation · Robustness

How well does the strategy avoid over-reliance on specific initial states, matchups, or situational advantages?

5.82%1,852 of 31,823 player-game observations reaching that player's second turn

1,488 wins4 draws · 360 losses

Early two-Prize turns were uncommon and often preceded wins. They occurred in 5.82% of the recorded player-game observations reaching the second turn. Of those 1,852 cases, 1,488 ended in wins (80.35%), but 360 ended in losses.

What this answers for robustness: it identifies a favorable situation across the public field. The archive summary does not compare those players’ results without the event or isolate our Lucario agent, so it cannot establish dependence on an early lead or prove that taking early Prizes caused the wins.

06 · Conclusion

Develop the deck and policy together.

We built Lucario and Dragapult agents with structured card information, PPO self-play and deck specialization. The results highlight these lessons:

  • Both specialists still outscored the generalist after deck edits. Local tests covered six familiar deck types, with the same 60 cards given to both players. After changing four Trainer-card copies in each list, Lucario’s game score fell from 66.93% to 63.09%, and Dragapult’s from 61.91% to 59.24%. Both stayed above the 50% even-result mark;
  • Overall performance can hide a difficult matchup. Lucario won only about one in four public games against Slowking / Kangaskhan, the leading-submission archetype of 4 of the official top 23 teams. The cause of this weakness and an effective remedy remain untested.
  • Public results differed by starting conditions. Both agents won more often going first on Kaggle. Lucario won 58.84% with opening Fighting Energy or Fighting Gong versus 47.41% with neither (+11.43 percentage points). Dragapult won 67.59% with opening Budew + Bench starter + evolution support versus 56.33% without the full combination (+11.26 points), though its adjusted association remained uncertain.

We are exploring joint deck–policy optimization at the archetype level, alongside search under partial observability.

08 · References

References.

  1. 01Kaelbling, Littman and Cassandra (1998) — Planning and Acting in Partially Observable Stochastic Domains
  2. 02Silver and Veness (2010) — Monte-Carlo Planning in Large POMDPs
  3. 03
  4. 04García-Sánchez et al. (2018) — Automated Playtesting in Collectible Card Games Using Evolutionary Algorithms: A Case Study in Hearthstone
  5. 05Yang et al. (2022) — PerfectDou: Dominating DouDizhu with Perfect Information Distillation