Expected Goals (xG) Explained: What It Actually Measures
An xG of 0.40 is not a claim that a particular player had a 40 per cent chance of scoring a particular shot. It is a claim about history. It says that in the database the model was trained on, shots taken from roughly that position, with that body part, from that kind of build-up and under that much pressure went in about 40 times in every hundred. StatsBomb put it more bluntly in 2017: an xG value "doesn't actually say much about this particular shot we are discussing right now. It's more like 'in the past, this has happened.'"
That distinction sounds pedantic. It is the entire argument. Nearly every complaint about expected goals - that it measures luck, that it decides who deserved to win, that it must be broken because a striker keeps beating it - comes from reading a historical frequency as a prophecy about one kick. What follows is what the number is, how it gets built, what it can carry, and the places where it quietly falls apart.
A number between zero and one
Opta, whose model most broadcasters use, defines it this way: expected goals "measures the quality of a chance by calculating the likelihood that it will be scored by using information on similar shots in the past." The scale runs from zero to one. Zero is a chance that could not be scored. One is a chance a player would be expected to convert every single time. Everything interesting happens in between.
Wyscout calls it "a predictive ML model used to assess the likelihood of scoring for every shot made in the game." American Soccer Analysis reaches for basketball and compares it to field-goal percentage, with the crucial difference that xG is tied to the location and the circumstances of the attempt rather than to the identity of the man taking it. That is deliberate, and it is why the number is useful at all.
The idea is older than the graphics package. Vic Barnett and Sarah Hilditch used the term "expected goals" in a 1993 paper on artificial pitch surfaces in English football. Richard Pollard and Charles Reep built the recognisable ancestor of the modern model in 1997, weighting shots by distance, angle, whether the ball was headed and how much defensive pressure there was. Jake Ensum, Pollard and Samuel Taylor ran a logistic regression on 930 shots from the 2002 World Cup. Sam Green's April 2012 post for OptaPro is what pushed the metric into the mainstream of football analytics, and in August 2017 Match of the Day announced it would put expected goals on television. None of those steps invented it on its own.
What the model looks at before it decides
Opta's current model is an XGBoost gradient-boosting machine trained on nearly one million shots drawn from 40 competitions across the 2018-19 to 2021-22 seasons. It evaluates more than 20 variables describing the situation up to the exact moment of contact: distance to goal, angle to goal, where the goalkeeper is standing, how clear the view of the goal mouth is, where the other players are, how much pressure the defenders are applying, the type of shot - which foot, volley, header, one-on-one - the pattern of play that produced it, whether open play, fast break, free kick, corner or throw-in, and the previous action or type of assist. For a sense of scale, the 2022-23 Premier League alone contained 9,609 shots, and the top five European leagues that season produced 45,764.
Other providers make different choices, and the differences are not cosmetic. Wyscout's public glossary lists exactly eight parameters, one of which is a human tagger's assessment of how dangerous the shot was. Understat uses neural networks trained on more than 100,000 shots with over ten parameters each. American Soccer Analysis publishes its logistic regression openly, coefficients included: a cross-assisted shot carries a penalty of -0.380 while a through-ball-assisted shot gets +0.909, which is the model's way of saying that a ball arriving across your body from the flank is much harder to finish than one played through the line. Hudl StatsBomb has the richest inputs of the lot: a computer-vision freeze frame recording the position of every attacker and defender visible in camera frame plus the goalkeeper, and the height of the ball at the moment of contact, because a ball at knee height is easier to strike cleanly than one at head height or rolling along the ground. StatsBomb reckons an open-goal flag alone can move a chance from 0.10 to 0.85.
Rough magnitudes, offered as illustration rather than as official figures: a shot from the edge of the six-yard box is often worth more than 0.35; a strike from 25 yards is usually below 0.03, and a speculative effort from outside the box around 0.02. Coaches' Voice has worked through a Mohamed Salah attempt valued at 0.08 and a Phil Foden goal at 0.46. Headers convert at lower rates than foot shots from the same spot. Weak foot converts lower than strong foot. Nothing on that list will surprise anyone who has played the game, which is a point in the model's favour rather than against it.
Penalties are the exception. Because the conditions are standardised - twelve yards, no defenders, no pressure worth modelling - providers simply assign a constant. That constant is about 0.76 to 0.79 depending on who is counting, reflecting a top-level conversion rate somewhere in the middle-to-high seventies. It distorts short-run totals badly: three penalties in a month adds well over two xG to a striker's account without him having beaten a defender. This is precisely why non-penalty expected goals, npxG, exists as a separate column, and why it is the column you want when comparing players.
The day the model changed its mind about a shot already taken
In March 2022 Opta shipped a new version of its model. It removed "big chance" as an input - a subjective tag applied by human analysts watching the match, long criticised for being both subjective and contaminated by whether the shot went in - and added goalkeeper position, defensive pressure and clarity of the shot.
The before-and-after cases are startling. A Richarlison chance against Wolves fell from 0.80 xG to roughly 0.20, purely because the new model could see where Jose Sa was standing. Roberto Firmino's against Watford rose from 0.81 to 0.96. James Maddison against Brentford and Patson Daka against Newcastle both jumped from about 0.47 to over 0.87 once goalkeeper position and the absence of defensive pressure were accounted for. The most extreme case reported was a Pedro Neto shot against Leicester that went from 0.035 to 0.656, nearly twenty times the value.
Nothing happened to those shots. They had been taken, saved or scored months or years earlier. What changed was the model's opinion about the category they belong to. That is the cleanest available demonstration of what an xG figure is: not a measurement of an event, but a classification of it.
It also retires the oldest criticism of the metric. "The model doesn't know where the keeper is" was true of early event-data models, has not been true of Opta's since March 2022, and was never true of StatsBomb's freeze frames. Big chance still exists in Opta's data, defined as a situation where a player should reasonably be expected to score, usually a one-on-one or a close-range effort with a clear path to goal and low to moderate pressure. It is a stat now, not a model input. Opta also runs an entirely separate model for women's football, trained on historical women's shots and in use since the 2023 World Cup, in which distance from goal carries more weight.
What a match xG line does not tell you
Opta is explicit that expected goals measures the quality of chances and "not the expected outcome of the game." The reason is arithmetic, and Jonas Lindstrom has the sharpest illustration of it.
Take two teams. Team Coin has four shots, each a 50 per cent chance. Team Die has twelve shots, each a one-in-six chance. Both finish the match on exactly 2.0 xG. They are not equally likely to have won. Team Coin's goal distribution is symmetrical around two. Team Die's is skewed to the right, which means the probability of it scoring few goals is larger than the probability of it scoring more. Over many replays, Team Coin wins more often. Twelve half-chances is not the same product as four good ones, and the xG total cannot tell them apart.
This is where the number stops and something else has to start. To get from an xG line to a result you have to simulate: either treat every shot as its own weighted coin and flip the match thousands of times, or feed each side's total into a scoring-rate model and do the same. That is the machinery behind expected points, where xPTS is three times the probability of a win plus one times the probability of a draw, and it explains a detail that puzzles people the first time they notice it - two teams' xPTS for a match do not add up to three, whereas real league points always do. If you want the mechanics of how thousands of simulated seasons become a title probability, our companion piece on how supercomputer predictions work covers exactly that, and the league simulator on this site runs the same idea over any table you care to build. Expected goals is the input to that process, not a substitute for it.
Two further warnings about single matches. Coaches' Voice puts it plainly: over one game expected goals "can give a skewed view." A side can dominate with three xG and still lose 3-0 to an opponent that scored early and then defended a lead, and neither the shot count nor the xG line is wrong about that. And fractions of goals cannot be scored. A total of 1.0 xG does not entitle anyone to a goal; it is an average across many imagined replays of the same set of chances, and plenty of those replays finish goalless.
How long before an xG number means anything
The honest answer is longer than one match, shorter than a season, and it depends. American Soccer Analysis ran the most careful test of this in July 2022, across 14,608 games in the top five European leagues from 2017-18 onward and 2,516 MLS games from 2018, with the 2020 pandemic seasons excluded. Their headline finding: "xG ratio is the best predictor of future points at any time during a season."
The detail matters more than the headline. xG ratio only overtakes plain goals ratio as a predictor after about eight games in the top five European leagues and about seven in MLS. In the original 2015 study, run on 2012-14 data, that threshold was four games. The gap between expected goals and simpler measures has narrowed since: shots-on-target ratio and total-shots ratio now sit much closer to xG than they used to. Anyone still repeating the 2015-era line that xG dominates every other metric is quoting an older sport. The top five European leagues, incidentally, proved far more predictable than MLS on every measure tested.
David Sumpter's practical guide is the most useful thing in the literature for a normal reader. Over one or two matches, expected goals tells you almost nothing beyond the scoreline. Over three to six, it is worth something only if the pattern is consistent. Between seven and sixteen matches is the sweet spot, the window in which a contradiction between a team's xG and its actual goals is genuinely worth investigating. Past sixteen matches, actual goals become the more reliable number and xG matters less for judging the season. Rolling averages over roughly ten games predict best.
Sumpter also produced the most humbling result in the field. A model built purely on a human being watching football and judging whether something was a big chance achieved an R-squared of 0.159, matching a complex multi-parameter xG model. His conclusion is worth keeping close: "expected goals are not football's magical equation." Watching matches still counts for something.
Beating the model, and drifting back towards it
Players and teams outscore their expected goals all the time, and there is nothing mysterious about it. Manchester City scored 99 league goals from about 91 xG in 2021-22. Erling Haaland, across a Bundesliga and Champions League spell, is reported to have scored 74 non-penalty goals from 57.76 npxG, roughly 28 per cent better than an average finisher would have managed with the same chances - the strongest evidence anyone had that elite finishing is a real, repeatable skill. Then in 2023-24 the same player underperformed by about 1.3 goals. Nothing about him had got worse.
That is regression to the mean, and it is routinely misread as a prediction of decline or a punishment for overachieving. It is neither. It is what happens when a short run of weighted coin flips keeps running. The persistence figures, reported from secondary analysis, make the point concretely: the year-to-year correlation for goals minus xG is around 0.12, which is close to nothing, while a measure of shot selection - shots over expected, relative to expected shots - comes in around 0.63. Getting into the 0.8 xG position again and again is the durable skill. Converting the 0.8 chance at a higher rate than everyone else, season after season, is far harder to prove.
Leicester City's 2015-16 title is the case study people reach for and usually get wrong. They won the league with 23 wins, 12 draws and 3 defeats, but sat around fourth on expected goal difference; on that measure the top three would have been Arsenal, Manchester City and Tottenham. They also posted the season's highest shot conversion rate at 13.0 per cent. Whether you call that finishing, luck, or a tactical system built to generate the kind of chances the model undervalues, the honest answer is that a single season's gap between goals and xG cannot separate those explanations.
That is the deepest problem with reading the gap at all. Goals minus xG blends finishing skill, goalkeeping at the other end, genuine luck and plain model error into one undifferentiated number. Published research on biases in expected goals models has shown that model bias contaminates estimates of finishing ability directly. The gap is a question, not an answer.
Which is why the players who dismiss the metric are usually half right. Morgan Rogers, then at Aston Villa, told Jamie Redknapp on Sky Sports in January 2026: "I think xG is a whole load of nonsense." He had seven goals from 4.21 xG at the time. "It feels like a crime - I'm scoring and xG is putting me down." He was factually correct that he had outscored the model. He was wrong about what the model had claimed: it never predicted he would score 4.21 goals, it said that shots like his had historically been scored about that often. Sean Dyche found the better register after Everton lost 1-0 at home to Fulham in August 2023 having created 2.73 xG to Fulham's 1.50. "One of our analysts said our xG, which I'm not that big a believer in but it's still a reference point, was around three, which is high in the Premier League." And then the sane part: "You don't examine a whole performance on xG... it's another marker, it gives you an idea."
Where xG breaks down, and the numbers built to patch it
Rebounds are the flaw a reader can catch for themselves on any given weekend. Shots are modelled as independent events, so a goalmouth scramble - save, rebound, block, rebound - accumulates expected goals shot by shot. Wyscout admits it in its own glossary: a sequence of shots in short succession "could theoretically yield an xG value of > 1," in a possession where only one goal was ever physically available. One published fix is to discount each follow-up by the probability it arises at all, so a rebound worth 0.8 that only occurs a quarter of the time, because the chance before it usually goes in, contributes 0.8 multiplied by 0.25, or 0.2. Not every provider applies a correction of that kind, and you should not assume a possession's xG is capped at 1.0.
Expected goals only scores shots that were actually taken. The most dangerous passage of a match - the cutback nobody attacked, the three-on-one that ended in a heavy touch - scores exactly zero. That blind spot is why non-shot xG and possession-value models exist: for each possession, take the xG of the shot if there was one, and otherwise the best xG that was available. A team whose non-shot xG runs far ahead of its xG is arriving in good positions and refusing to shoot from them.
Team totals also compare unlike things. A low block deliberately compresses the high-value zones and forces attempts from outside or from tight angles, worth well under 0.05 each. A low expected goals against can therefore mean "defends deep" rather than "defends well." A counter-attacking side takes far fewer shots but takes them against a disorganised, outnumbered defence, so each one is worth more. Comparing raw xG or xGD between a possession team and a low block is comparing two different products.
Standard models are shooter-agnostic on purpose. They answer how often a chance like this is scored, never how often this particular player scores it. Building the finisher into the model would make it impossible to use the model to judge the finisher. And with thirty-odd shots in a season and a binary outcome each time, one player's goals-minus-xG carries enormous uncertainty. Treat it as a flag, not a verdict.
A family of adjacent numbers exists to cover these gaps, and they should not be conflated. npxG strips out penalties. xGA is the expected goals of the shots a team conceded, and xGD is xG minus xGA. Expected assists measures the likelihood that a completed pass becomes a goal assist, considering the type of pass, its length and its end point. And expected goals on target - xGOT, or post-shot xG - is the one people confuse with xG most often. Expected goals covers everything up to the moment of contact; xGOT covers everything after it, combining the underlying chance quality with where in the goal mouth the ball actually finished. It exists only for shots on target, so a miss has an xG and no xGOT. Harry Kane in 2019-20 had 10.9 xG and 14.5 xGOT, meaning placement alone was worth about three and a half goals. Dean Henderson faced 39.4 xGOT that same season and conceded 32, so he prevented around seven. The difference between the two is sometimes published as Shooting Goals Added.
Who publishes xG, and why no two of them agree
Every provider trains its own model on its own event data, with its own definitions, its own variables and its own weights. The consequence, in the flat summary the literature settles on, is that values from different providers are "not necessarily directly comparable." The structural reasons are already above: Opta's XGBoost over nearly a million shots and more than twenty variables, having dropped its subjective big-chance tag in 2022; StatsBomb's freeze frames of every visible player plus ball height; Wyscout's eight parameters including a human's judgement of danger; Understat's neural network. Even the penalty constant differs - 0.76, 0.78 or 0.79 depending on the provider. Divergence on the same shot is commonly reported at around 0.05 to 0.10 xG.
The effect at season level is not trivial. Son Heung-min scored 23 Premier League goals in 2021-22 and shared the Golden Boot with Mohamed Salah, the first Asian player to win it, and none of his goals came from the penalty spot, which makes him a natural case study in finishing. His expected goals for that season is published as 13.95 in one place and around 16.0 in another. Same 23 goals, same shots, two meaningfully different verdicts on how much of it was finishing and how much was chance quality. Pick a provider and stay with it.
As for where the numbers live: Opta and Stats Perform are the enterprise standard behind most clubs and broadcasters. Hudl StatsBomb sells the richest data and also publishes a free open data set covering women's competitions, men's internationals and some historical seasons. Wyscout, also Hudl, is the scouting industry's video standard with the largest match library. Understat is the main genuinely free source of high-quality shot-level xG and npxG, covering the top five European leagues plus the Russian Premier League back to 2014-15. For casual lookups, FotMob carries live xG, and Squawka, SofaScore and StatMuse publish it too.
One piece of standard advice is now out of date, and most articles have not caught up. On 20 January 2026 Stats Perform terminated FBref's data agreement and required the immediate removal of every advanced statistic it supplied - expected goals, progressive passes, shot-creating actions, all of it. Sports Reference's president, Sean Forman, announced the change; Stats Perform said FBref had violated the terms of the agreement and did not specify how. The site kept scores, goals, assists and appearances, and the advanced archive stopped updating. The termination came eight days after Stats Perform was named FIFA's exclusive betting-data and streaming-rights distributor for the 2026 World Cup, which fuelled a good deal of speculation about the commercial motive. FBref said it was "especially upset by the massive step back this creates in access for women's" advanced data. Unless that reverses, Understat is where a reader without a budget should look.
Which is the last thing worth understanding about expected goals. It is not a public fact about football, like the score or the attendance. It is a product, built by a handful of companies out of proprietary data, using methods they can revise - and have revised - with real consequences for how a shot taken years ago is remembered. Read it as what it is: a well-founded estimate of how often chances like these have gone in, and the best available starting point for the question of what happens next. What happens next still has to be simulated, and it still has to be watched.
Simulate a season and watch chances turn into results
An xG model does one thing: it takes a chance, weighs everything it knows about it, and returns the probability that the ball ends up in the net. The simulator on futsimulator.com runs on the same premise, one level up. It takes what each side is capable of, turns that into a probability for every fixture, and then resolves it into something that actually happened: a scoreline, three points or one, a table that sorts itself. You are never being told what deserved to happen. You are watching a chance become a result, again and again, until the pattern underneath shows through.
What that gives you is the thing a single match xG line can never show you: the spread. Run the same league twice and the side you had down as champions can finish fourth with nothing about the squad changed. Run it often enough and the good teams win most of the time, which is what the numbers were saying all along. A full season is long enough for the model to be right; one Saturday afternoon is not. Seeing that with your own table in front of you lands harder than reading it in an article.
"xG is not a verdict on what should have happened. It is a count of how often that shot goes in."
- Simulate a full league season and let the table sort itself, then compare the finishing order with the one you would have written down before the first matchday.
- Run the same season again, and then a third time, and count how many points separate the best run from the worst with identical squads on both.
- Simulate a single match on repeat: same fixture, a different scoreline nearly every time, which is what a chance worth 0.30 looks like from the inside.
- Build a knockout bracket and play it through to the final, where a tie is a tiny sample and the better side goes out far more often than the table says it should.
"Play enough matches and the luck cancels out. That is the only thing xG has ever asked you to do."
