Third weekend of a season, and the broadcast graphic goes up. A striker converting better than one chance in two. A goalkeeper with a save percentage in the mid-nineties. A team described as being on pace for a record points total. Every figure on that screen is arithmetically correct, and not one of them supports the sentence being read over it.
The mistake is not bad arithmetic. It is a category error: treating a number built from three matches as though it were the same kind of object as a number built from thirty-eight. They look identical on screen. They behave nothing alike, and the difference is not a matter of degree.
What follows is the triage I use before quoting anything in September. Some numbers are usable on the opening weekend. Some need a dozen matches or more before they carry any information at all. And a few never settle enough to be quoted about an individual, in any month.
Triage table: what settles, and roughly when
| Settles | Type of metric | Examples | Why |
|---|---|---|---|
| Almost immediately | High-count volume | Passes attempted, shot volume for and against, possession share, distance covered, corners, minutes distribution | Hundreds of events per match, largely chosen by the team rather than granted by luck |
| Within a handful of matches | Structural and profile | Shot locations, pressing height, set-piece routines, formation, who takes what | Repeated deliberate choices, visible whether or not they produce goals |
| Ten to fifteen matches | Team-level rates on frequent events | Duel success, pass completion by zone, shots on target share | Denominator grows fast enough to outrun the noise, but not instantly |
| Half a season or more | Rates on rare events | Conversion rate, save percentage, penalty record, clean-sheet counts | Denominators in single or low double digits, where luck dominates |
| Rarely, at individual level | Ratios of ratios and low-minute per-90 figures | Goals per shot on target, anything per 90 built on a few hundred minutes | Noise in both numerator and denominator, compounding |
The two forces that make early season stats lie
The first force is arithmetic. A percentage built on a small denominator can sit far from its true value purely by chance, and the smaller the denominator the wider that range gets. Four attempts and three conversions produce a headline rate of seventy-five per cent, which tells you almost nothing about the fifth attempt.
The second force is selection, and it is the one people miss. The reason a number reaches a graphic is almost always that it is extreme. Nobody builds a chyron around a striker converting at exactly the league average. So the numbers you see in week three have been filtered for being unusual, and being unusual in a small sample is mostly a symptom of luck rather than a sign of quality.
Put the two together and you get the pattern anybody who follows early season stats will recognise: the most spectacular figures of September are, as a group, the ones that will fall furthest by Christmas. Not because the players get worse, but because they were never really up there.
Volume settles first, efficiency settles last
Sort any metric by how many events feed it per match and you have most of the answer. A team attempts hundreds of passes and a couple of dozen shots every week. It scores a small handful of goals, concedes fewer, and takes a penalty perhaps once a month.
Volume metrics also have a second advantage: they are largely chosen. A side that presses high does it every week because that is the plan, not because the ball bounced kindly. That makes shot volume, pressing height and possession share readable almost immediately, while the goals those shots produce stay unreadable for months.
The practical consequence is that the interesting September question is never how many a team scored. It is how many shots it took, from where, and how many it allowed. Those three answers are available in week two and they are still true in April.
The ordering holds across codes, which is a good sign that it is structural rather than a quirk of one sport. In basketball, shot attempts, rebound opportunities and pace settle within a handful of games, while three-point percentage stays unstable long enough to ruin plenty of early-season narratives. In hockey, shot attempt shares are readable early and shooting percentage is not. Different sports, same rule: count the events feeding the number.
The denominator nobody checks

Most bad September numbers share one structural fault: the denominator is either small or invisible. A per-90 figure calculated on a couple of hundred minutes is a rate built from a fragment. A conversion percentage on six attempts is a coin-flip wearing a suit.
There is a fast diagnostic. Any rate sitting at exactly zero or exactly one hundred per cent is telling you about its denominator rather than about the player, because those values only survive when the sample is tiny. The same goes for suspiciously round percentages: fifty, thirty-three, twenty-five. Those are small integers in disguise.
The habit worth building is to refuse to read any rate that does not arrive with its denominator attached. Five goals from twenty-two shots is information. Twenty-three per cent conversion is the same fact with the useful half removed, and it is the half that tells you whether to believe it.
There is a quieter denominator problem too, and it is definitional. Two data providers can count the same match and disagree about what a tackle, a key pass or a duel actually was, because each works from its own written definition. Over a full season those differences mostly wash out into a consistent house style. Over three matches they can be the entire gap between two players, which is why mixing sources in September produces comparisons that fall apart the moment somebody checks where each figure came from.
Why a model number moves before the real number does
Shot-quality models, of which expected goals is the best known family, get treated as either magic or nonsense. They are neither, and their real advantage in September is unglamorous: they use every shot, not only the ones that went in.
That difference in effective sample size is why a team’s shot-quality total per match starts to look stable while its goal total is still jumping around. The model is counting twenty or so events per match where the scoreline counts two or three. It is not smarter about football; it simply has more to work with.
Two caveats keep it honest. Different providers build these models differently, so figures from two sources are not interchangeable, especially over a few matches. And a model trained on how shots usually behave will misread a team that is deliberately doing something unusual. Performance analytics is at its strongest describing patterns and at its weakest judging a single unusual case.
September’s table is partly a fixture list

After four or five rounds, no two teams have played comparable opposition. One side has had three home matches against sides that finished in the bottom third; another has been away twice against sides with continental commitments. The points column records both as the same currency.
Schedule effects fade because the fixture list eventually balances, but in the first month they can be larger than any real difference in quality between mid-table clubs. Adjusting for them is standard practice in matchday analysis and almost never present in the on-screen graphic.
There is a fixture-congestion layer as well. Squads in continental competition rotate more, travel more and play tired more often, and their early numbers reflect a calendar rather than a level. Reading the table in September without reading the fixture list beside it is guesswork with extra steps.
The sentence that breaks: on pace for
Linear extrapolation is the most confident-sounding mistake in sport. On pace for assumes three things at once: that the observed rate is the true rate, that nothing about the schedule changes, and that the player or team keeps the same role for the rest of the season.
All three assumptions fail routinely. Rates regress. Schedules get harder or easier. Roles change with injuries, form and transfer windows, and a striker averaging a goal a game across three fixtures against promoted sides is not carrying that rate into a run against the top four.
The repair is easy and costs one clause. Instead of on pace for, write what has actually happened and what would have to hold. Has scored in each of the first four matches, a rate that no player in the division sustained last season. That sentence keeps the drama and drops the false precision.
How analysts find the point where a number settles
The standard method is simpler than its reputation. Split a large historical sample in half at random, odd matches against even matches, and check how strongly a team’s value in one half predicts its value in the other. Repeat at different sample sizes and find the point where that relationship becomes strong enough to be worth acting on.
The reason this works is that a metric made mostly of skill will agree with itself across two random halves, while a metric made mostly of luck will not. That is also why the answer is different for every metric and every level of competition, and why anybody quoting one universal threshold for all numbers is overselling.
The underlying behaviour has a name outside sport as well. Regression toward the mean is the general tendency of extreme measurements to be followed by less extreme ones, and it applies to shooting percentages exactly as it applies to exam marks and blood pressure readings.
Numbers you can quote in week three
- Shot volume, for and against. Dozens of events per match, and directly tied to how a side sets up.
- Shot locations. Where a team shoots from is a choice, and it shows up almost immediately.
- Possession and pass volume. Hundreds of events per match, stable within a couple of fixtures.
- Pressing height and defensive line position. Repeated deliberate behaviour, visible on any broadcast.
- Set-piece volume and who takes them. Structural, countable and unaffected by finishing luck.
- Minutes distribution and rotation patterns. Facts about selection, not estimates.
- Distance and high-speed running from tracking data. Very high event counts, and therefore quick to settle.
Numbers to leave alone until November
- Conversion rate and save percentage. Both are rates on rare events, and both are the classic September mirage.
- Points per game as a measure of quality. Contaminated by schedule before the fixture list balances.
- Goal difference read as dominance. One heavy win distorts it for months.
- Any per-90 figure built on a few hundred minutes. The denominator is doing all the work.
- Clean sheets and penalty records. Countable, but far too rare to separate skill from sequence.
- Form streaks. A run of results selected after the fact is not evidence, it is a highlight of the schedule.
Writing the honest version of the same sentence
None of this requires dropping numbers from a match report. It requires three habits. Quote the count with the rate, so the reader can weigh it. Name the baseline, because a percentage without a league average is a decoration. And state the condition, meaning what would have to continue for the figure to survive.
Compare two versions of the same claim. First: he is converting at thirty-eight per cent, the best in the league. Second: he has five goals from thirteen shots, a rate roughly three times the divisional average and one that nobody sustained over a full season last year. The second is longer by a clause and honest by a mile.
The same discipline is what separates a durable argument from a disposable one when advanced metrics get used to rank individuals. A ranking that changes every time three more matches are played was never measuring what it claimed to measure.
Frequently Asked Questions
Is fifteen matches a real threshold or a rule of thumb?
A rule of thumb, and deliberately a rough one. In a weekly-fixture sport it is a reasonable landing zone for team-level rates on moderately frequent events. Volume metrics arrive far earlier, rates on rare events arrive much later, and each metric has its own answer that has to be measured rather than assumed.
Can combining several metrics fix a small sample?
Partly. Averaging several noisy indicators does cancel some of the noise, which is why composite ratings look steadier than their components. The cost is transparency: a composite that moves gives you no idea which underlying number moved, and a small sample inside the composite is still a small sample.
What about a genuinely dominant individual performance?
Watch it and describe it. What you saw is evidence about that match, and structural observations hold up fine: he played wider, he took the free kicks, the team pressed higher. The projection is the part that fails, not the observation.
Should last season’s numbers be mixed in?
That is the standard repair. Blending a small current sample with a longer historical baseline gives a far better estimate than either alone. The obvious caveat is change: a new role, a new system or a different level of competition can make the older data misleading in a different direction.
Why do broadcasters keep using September rates?
Because they are available, instantly comprehensible and dramatic, and a graphic has about six seconds to land. The problem is rarely the number itself. It is the sentence wrapped around it, and that sentence is free to fix.
Keep a note of the three most striking figures you see this weekend, then look them up again in mid-November. The exercise takes two minutes and it will change how you read a stat graphic for good.
