How Our Tennis Model Works: Elo, Serve Stats, Monte Carlo
Tennis is one of the cleanest sports to model — one player against another, no teammates diluting the signal, decades of match data freely available. That simplicity is deceptive, though. Getting a model to genuinely add information beyond what a sharp market already prices is hard, and getting a model to be honest about the markets where it can't is harder still. This post walks through exactly how our ATP and WTA models are built, what feeds them, and — in the one place it matters most — where we deliberately chose to bet less rather than pretend to know more.
Two tours, two models
We train separate models for ATP and WTA. Men's and women's tennis have genuinely different dynamics — serve dominance, match length, rally patterns — and forcing one model to cover both tends to average away real signal. Both are single XGBoost classifiers trained on the match-winner (h2h) market, using ~23,000 ATP and ~21,000 WTA matches from 2018–2026, sourced from Jeff Sackmann's open historical dataset.
Skip the hand-calculation.
Get real value bets flagged for you — 7-day free trialThe resulting Brier scores — a standard measure of probabilistic accuracy where lower is better and 0.25 is what a coin flip scores — are 0.1629 for ATP and 0.1547 for WTA. Both sit well below the coin-flip baseline, which is the first sanity check any prediction model has to clear before it's worth building anything on top of it.
One deliberate design choice: the two tours don't even use the same number of features. ATP runs on 18 features, WTA on 28. That's not an oversight — it's the result of per-tour feature pruning validated out-of-sample. A feature that sharpens the WTA model can be pure noise on ATP, and vice versa. The clearest example: WTA drops overall Elo entirely from its feature set and relies instead on surface-specific Elo — a distinction worth its own section.
Elo, and why surface matters
If you're unfamiliar with the mechanics, our Elo ratings explained piece covers the general concept — a self-correcting rating system that moves up after wins against strong opponents and down after losses to weaker ones. Tennis adds one wrinkle worth knowing about: a player's overall Elo is a blend of results across clay, grass, and hard courts, but those surfaces play so differently that a single number can hide a lot.
So alongside overall Elo, we compute surface-specific Elo for clay, grass, and hard courts separately. A player who's a clay-court specialist but mediocre on grass will show that split clearly in surface Elo, even when their overall number looks unremarkable. Because a lot of players simply haven't played many matches on one surface — grass, especially, gets a tiny fraction of the calendar — surface Elo is shrunk toward the overall rating when the sample on that surface is thin. This keeps a specialist's grass rating from being wildly overconfident off three matches.
This is exactly why the WTA model can afford to drop overall Elo and lean on surface Elo instead: once you have a reliable surface-specific number, the blended overall rating adds little on top of it — for WTA, our validation showed it added nothing but noise.
Serve, return, fatigue: the feature set
Beyond Elo, the models lean on three families of signal:
- Serve and return stats — first-serve win percentage and return points won, both proven predictors of match outcomes at the professional level.
- Fatigue and form windows — days since a player's last match, matches played in the last 7 days, and recent win rates.
- Elo — overall and surface-specific, as above.
One of these deserves a callout most bettors wouldn't guess: days since last match is one of the single strongest features in the entire model. Remove it and Brier gets worse by 0.015 — a large move for one feature in a model this size. Fatigue, or the lack of it, genuinely moves match outcomes on tour.
We also tested head-to-head record as a feature — the classic "Player A always beats Player B" intuition — and dropped it. On out-of-sample validation it added nothing once Elo and form were already in the model; head-to-head record mostly just re-states relative skill that Elo already captures, with a lot of extra noise from small sample sizes.
Point-in-time discipline: no lookahead, ever
Every serve and return statistic fed into the model is calculated strictly from the prior season — never from stats that include the match being predicted, or any match after it. This sounds obvious, but it's the single most common way a sports model quietly cheats itself: if training data ever lets a feature "see" a full season's aggregate stats on a match that happened mid-season, the model learns a pattern that simply won't exist when predicting a real, not-yet-played match. Serve skill is stable enough year over year that using the prior season's numbers loses almost nothing — while eliminating that leak entirely.
The same discipline applies to Elo, which is why it has to stay current. Our Elo ratings refresh nightly. We learned this the hard way: a stale Elo input — even a few weeks out of date — quietly costs several points of accuracy, even when every other part of the model is working exactly as designed. A model is only as sharp as the numbers going into it on the day it's asked a question.
Monte Carlo totals — and why the model learned to say no
Match-winner isn't the only market tennis offers. Total games (Over/Under) is priced differently — not by the XGBoost classifier, but by a Monte Carlo simulation: we simulate the match game-by-game using both players' estimated serve-hold probabilities, thousands of times, and read off the distribution of total games played.
Here's the honest part. When we audited the simulator's Over/Under probability against actual settled results, the numbers were not good. The correlation between the model's predicted probability and what actually happened was essentially zero (≈0.02) — meaning the simulator's confidence carried almost no real discriminating information — and it carried a systematic bias toward Over. A model that's both undiscriminating and biased is exactly the kind of thing that looks fine in a demo and loses money in production.
Rather than patch over that with a display filter, we fixed the actual output. We fitted a per-tour Platt calibration on the totals probability — a standard statistical technique that maps a model's raw output onto its true, honest reliability. Because the simulator had almost no real discriminating power to begin with, the calibration honestly does what an honest calibration should: it collapses the probability toward the base rate. The practical effect is that the model now largely abstains from totals bets rather than manufacturing a false edge. Only the cases where a soft book's price is genuinely mispriced against Pinnacle's de-vigged line — read de-vigging explained for how we extract that fair line — clear our EV floor and get surfaced as a bet.
Match-winner betting was not affected by this — that market is priced by the XGBoost classifier described above, which we validated separately and which discriminates well. This calibration episode is specific to the Monte Carlo totals simulator. For the fuller story of how the models have evolved, see our tennis model v3 upgrade.
From probability to bet
A model probability alone isn't a bet — it only becomes one when it's compared against a market price. We anchor every tennis prediction against de-vigged Pinnacle odds, the sharpest publicly accessible tennis market, and treat that de-vigged number as our best available estimate of true probability. Our own model has to disagree meaningfully with that anchor — and a soft book has to be offering odds better than what the de-vigged Pinnacle price justifies — before anything gets flagged.
One honest coverage caveat: The Odds API, our odds data source, only carries roughly 36 named tournaments — Grand Slams, Masters events, and a handful of 500-level tournaments. Many smaller 250-level events simply have no Pinnacle line available, and our rule here is absolute: no Pinnacle anchor, no bet. We won't estimate a fair line from thinner soft-book prices alone; if the sharp reference isn't there, we sit that match out entirely, regardless of how confident our own model might otherwise look.
When a genuine edge does clear the bar, we require a minimum 5% expected value, cap it at 18% (past that, it's almost always a data artefact rather than a real edge), and cap max odds at 3.50 — long-shot prices are the least reliably calibrated part of any market. We also run slightly tighter Pinnacle-ratio anchors on WTA than on ATP, a distinction that came directly out of validating against our own settled bet history rather than a theoretical assumption.
Stake sizing follows fractional Kelly at 25%, capped at 5% of bankroll per bet — a deliberately conservative fraction of the mathematically "optimal" Kelly stake, because a model's edge estimate is itself uncertain, and full Kelly punishes that uncertainty severely over a long betting history.
How we know it works — and where it doesn't
Brier scores like 0.1629 and 0.1547 tell you the model beats a coin flip, but they don't tell you whether it's actually finding profitable mispricings against a sharp market. For that we track closing line value (CLV) — did the price we bet at improve or worsen by the time the market closed? Read closing line value explained for the full mechanics. CLV is the metric we trust most for match-winner bets specifically, because that market behaves like a classic moneyline: a wide enough price spread and an efficient close make CLV a reliable long-run signal. For totals, we lean more on realized ROI and calibration checks than on CLV alone, because totals markets behave differently at the margins.
We publish every settled bet, win or lose, on our track record page — including the ones that didn't work out. You can also browse live model outputs and today's coverage on the tennis hub, and read more about the model itself on our model page.
None of this is a promise of profit. Variance dominates any small sample of bets, tennis markets can be thin outside the biggest events, and even a well-calibrated model will have losing stretches that mean nothing about whether the underlying edge is real. What we can promise is that every number you see — the model probability, the Pinnacle anchor, the expected value — is exactly what it claims to be, calculated the same honest way whether it produces a bet or an abstention.