How the numbers work
Where Pick Four's data comes from, how often it changes, and what each number counts
Where the data comes from
Every usage number says where it comes from: hover a source line or a usage bar for the full details. Regulation M-C has these sources now, and the format menu and the Meta page switch between them:
- Ladder games · 21–28 Sep: Showdown ladder games · gen9championsvgc2026regmc · 2026-09-21/2026-09-28 · 1500+ rating · 1,310 in sample · fetched 28 Sep 2026
- Tournaments · 12–28 Sep: Limitless tournaments · M-C · 2026-09-12/2026-09-28 · 4,567 in sample · fetched 28 Sep 2026
- Ladder games are rated games from Pokémon Showdown’s public replays: Bo1 games at 1500+ and Bo3 games at 1300+, from the start of the last complete week (Monday, UTC). Players’ names are never kept. Each side gets a player code, keyed with a secret that every import draws and throws away, so it only tells which games in one week one player played.
- Tournaments are completed events of the last four weeks with 16 or more players, from Limitless, and official Play! Pokémon events (Regionals, Internationals, Worlds) from Limitless VGC once their team lists are out: placings, records, open team lists and match results, never players’ names.
- Smogon stats are Smogon’s monthly ladder statistics at the 1760 cutoff. A regulation gets them once Smogon publishes its first month.
- The development fixture is a small sample used only by tests, or by a regulation with nothing live.
When it changes
The data refreshes every Monday at 06:17 UTC: last month’s Smogon stats, the ladder games of the week that just ended, and new tournaments. The refresh checks the data, runs the tests and builds the site before it goes live; if anything fails, nothing changes until it’s fixed.
Usage
- Usage is the share of teams with the Pokémon: two teams a ladder game, one a tournament entry.
- Brought and leads (ladder games only) are the share of its teams that sent it into battle, or sent it out first.
- Moves, items, abilities and natures come from open team sheets (every Bo3 ladder game) and tournament team lists, as the share of the sheets that show the Pokémon. Teammates are the share of its teams with each other Pokémon.
Win rates
- A win rate counts decided games or matches: ties, byes and matches nobody won are left out.
- Each player’s games with one team count once, as that player’s share of wins, so someone who replays a team forty times weighs as much as someone who played it once. A tournament entry is one player.
- The range after a rate is a 90% Wilson score interval over those players: with few of them it’s wide. A rate is green or red only when its whole range is above or below 50%.
- Fewer than 20 decided games or 5 players is a small sample. Below 20 games you see the won–lost record instead of a rate.
- Sorting by win rate ranks by the range’s low end, with small samples last, so a few lucky games can’t top a list.
- Teams in a tournament’s top cut won most of their matches, so compare win rates within one cut.
Worked examples (win-rate@1)
- 15–10 in 25 games by 25 players: 60.0%, range 43.7%–74.4%.
- 22–5 in 27 games, but 20–2 of it by one player: 81.5% over all games, 48.5% with each of its 6 players counted once (range 21.1%–76.8%).
- 3–0 has a range from 52.6%, above 60–40 (51.8%), but it’s a small sample, so it ranks after it.
Trends
- Trends split the ladder games into weeks (Monday to Sunday, UTC). A week’s usage is the share of its teams with the Pokémon, and its win rate counts each player’s games with one team once, as above; a week needs 20 decided games for a rate.
- Rising and falling compares the last two weeks with 100 or more games. The current week counts once it has that many, marked “so far”.
- A Pokémon is listed when its share moved 2 points or more and a two-proportion z-test puts the change beyond chance at 95% (|z| of 1.96 or more): a change that size, between samples that big, would happen by chance less than 5% of the time.
- A trend says what players are picking more or less often, not whether it wins: the win rate beside it does that.
Worked examples (trend@1)
- On 50 of 500 teams one week and 80 of 500 the next (10.0% to 16.0%, +6.0 points): z = 2.82, so it’s rising.
- 50 of 500 to 60 of 500 (+2.0 points): z = 1.01, within chance, so it isn’t listed.
Tier list
- The tier list takes every Pokémon on 1% or more of ladder or tournament teams and scores it out of 100 from three parts, each with its reason on the page.
- Usage (60%): its share of teams on a log scale, the most used at 100, so going from 1% to 5% counts for more than from 30% to 34%. The ladder’s and the tournaments’ (every entry) are averaged.
- Ladder results (20%): 50, plus 5 for each point its teams’ win rate’s 90% range sits above 50%, or minus for each point below. A range across 50% counts as even, so a small sample moves nothing. From 30 decided games.
- Tournament results (20%): the same for its top 8 conversion against the field’s rate, relative to it (a range 10% above the field’s adds 10). From 20 entries.
- A part without enough data leaves its weight to the others. Rising or falling on the ladder this week (see Trends) adds or takes 3. S from 74, A from 62, B from 53, C from 45, D below.
- One-on-one calcs (the meta score on Counters) aren’t counted: they underrate support Pokémon such as Indeedee and Whimsicott, whose value is what they do for a partner. The tiers describe the week’s games; a Pokémon in C can still be the right pick for your team.
Worked example (tier@1)
On 20% of ladder and 15% of tournament teams, when the most used are on 50% and 60%: usage 72.4. Its ladder teams won 54% with a range of 51–57%, 1 point clear of 50%: 55.0. It made the top 8 with 15% of its entries, range 11.3–19.4%, across the field’s 13.5%: 50.0. Score 64.5: A tier.
Sets and stat points
A Pokémon’s most common set is built field by field: moves, item, ability and spread each come from the first source that has them (the chosen source, then the regulation’s other live sources, then the previous regulation’s). Team sheets and team lists show natures but no stat points, so when no source has a spread the set takes the most common nature with stat points generated to match it. Each set says where its parts came from.
Estimates
Scores and simulations are estimates from simple models, labelled “Estimate · model” wherever they appear, each with the reasons behind it:
matchup@1: one Pokémon against another, from damage calcs and speed.meta-score@1andcounter@1: how a Pokémon does against the most used ones, weighted by usage.team-score@1,partner-score@2andcounter-team@1: team scores, build-around partners and the best team against the meta.team-archetype@2: which game plans a team fits.team-lab@3: the team builder’s recommended teams and the team of the week (see below).spread-optimiser@1: the KO finder’s spread optimiser.sim@3: the matchup simulator (see what it leaves out).
Checking the simulator
Each week the simulator plays the Bo3 ladder games whose players showed their open team sheets, as the simulator page would (Auto bring-4, usage-weighted decisions, 1,000 games each), with the first player’s team as yours. Sheets show no stat points, so each Pokémon has its usual spread for the sheet’s nature. Its estimates are then set against who won.
- The Brier score is the average of (estimate − result)², a win counting 1 and a loss 0: 0 is perfect, and always saying 50% scores 0.250. Skill is how much of that 0.250 the estimates save, so above 0 beats a coin flip.
- Real games are closer than simulated ones: the simulator knows the teams, not the players. “Edges held” is how strongly its edges held up in these games, and the simulator page shows each estimate pulled that far towards 50% (“in real games, about”). The factor is refitted each week; fitted on half the games, it’s scored on the other half.
- Reliability puts the games in buckets by estimate: when it says 60–80%, the side should win about that often. A bucket’s range is wide with few games.
- After a run, the simulator also shows what the ladder’s games say about that matchup (
ladder-evidence@1): the record of these two teams, of similar teams (4 or more of each side’s six) and of each Pokémon against each, and whether the estimate falls inside the first of these with enough games.
- Brier score
- 0.253
- always 50%: 0.250
- Skill
- −1.2%
- above 0 beats a coin flip
- Favourite won
- 55.8%
- of the games
- Edges held
- 45%
- pulled in: 0.247
- Games
- 265
- 1,000 simulated each
- Brier score 0.253 over 265 games: 0 is perfect, and always saying 50% scores 0.250.
- Skill −1.2%: always saying 50% would have done better.
- The side the estimate favoured won 55.8% of the games.
- On average it gave the side simulated as yours 49.5%; that side won 44.5%.
- In 3 of 5 buckets with games, the mean prediction is inside the 90% range of how often that bucket's side won.
- Its edges held up about 45% as strongly in real games. Pulled that far towards 50%, with the factor fitted on the other half of the games, the estimates score 0.247.
| Predicted | Games | Mean | Won (90% range) |
|---|---|---|---|
| 0–20% | 8 | 14.7% | 12.5% (3–41) |
| 20–40% | 72 | 31.8% | 38.9% (30–49) |
| 40–60% | 112 | 50.1% | 46.4% (39–54) |
| 60–80% | 60 | 66.8% | 50.0% (40–60) |
| 80–100% | 13 | 83.1% | 53.8% (33–74) |
Estimate check · sim@3 · sim-calibration@1 · checked 29 Sep 2026
Source: Showdown ladder games · gen9championsvgc2026regmc · 2026-09-21/2026-09-28 · 1300+ rating · 265 in sample · fetched 28 Sep 2026
Recommended teams
Each Monday the team lab builds complete teams from every Pokémon on at least 0.5% of the week’s ladder or tournament teams, each with its most common set, searching on the team score (team-score@1):
- Each game plan, fit first: a point for each point of its fit, so the team reaches the best fit it can, then the team score and 0.2 of a point for each point of its strength against the meta (how well its best answer to each of the 20 most used Pokémon does, weighted by usage). A plan that needs moves few players run (hazards, screens, Tailwind, Trick Room, redirection, Perish Song) can use a plan set: a Pokémon’s usual set with the plan’s move in place of its least used move, never Protect (an attack only from the side the set attacks from, physical or special), and the team says which sets it changed. Weather and terrain plans keep their real setters, and play styles keep the sets players run. Hazards count kinds: Stealth Rock and Spikes are two, two Stealth Rock setters are one. Each plan gets two teams: its best, and one built around its most dedicated setter (the fewest attacks, then the most bulk: Toxapex with Toxic Spikes for hazards). The simulator picks the better one.
- Counter-meta: 75% strength against the meta and 25% team score, with 1 point for each off-meta member (on under 10% of teams), up to 2: enough to pick an off-meta Pokémon over a common one that answers the meta about as well, not over a clearly stronger one.
- The team score alone: its best team, a check on the score: the simulator seldom agrees with it.
Every team keeps the item clause, the species clause and one Mega. Each is then played 1,000 times against each of the ladder’s 10 most used teams of the week, as the simulator’s gauntlet plays them (Auto bring-4 on both sides), and its win expectancy counts each of those teams by its players. A team is “ahead” in a matchup when the whole 95% range of its win rate is above 50%.
Then the simulator refines the best, since the team score doesn’t always see what wins: each member of the 3 best teams is swapped for its 2 best replacements, ranked half on the team score and half on how well the team then answers the Pokémon on the teams it loses to. Every swap is tried in quick games (200 games a matchup, a shorter pre-sim) and the most promising 4 are played in full; if one beats every team so far, a second round refines it. The best of all is the team of the week (team-lab@3). The same teams and data always give the same results, so the lab only runs again when something changed.
Damage numbers
Damage and stats aren’t estimates: every roll comes from the engine, which uses @smogon/calc’s Pokémon Champions mechanics with stat points, the game’s 4096-based modifiers and its rounding, and matches every test case exactly, all 16 rolls. Anything around them (who moves first after Tailwind, what a spread survives) is worked out from those numbers.