Methodology

This page explains how Proknoz generates its predictions. Transparency is a core principle — you should understand what the model does, what data it uses, and where it might be wrong.

Football — Dixon-Coles Model

Football predictions use a variant of the Dixon-Coles model, which extends the standard Poisson regression by correcting for the observed correlation between low-scoring outcomes (0-0, 1-0, 0-1, 1-1).

The model estimates attack and defence strength parameters for each team, along with a home advantage factor. These are fitted using historical match data with an exponential decay that weights recent matches more heavily than older ones. A regularisation term keeps these parameters anchored toward the population average, so a team only earns an extreme rating when the evidence genuinely supports it (see Calibration and Accuracy below).

At prediction time, four multiplicative adjustments are applied to the base expected goals (lambda) before computing match outcome probabilities:

From the adjusted expected goals, a Poisson score matrix is computed and used to derive probabilities for all standard markets: match result (1X2), over/under 2.5 goals, and both teams to score.

For international fixtures (World Cup, European Championship), the same model is trained on national-team results going back several years, including competitive matches and friendlies. Friendlies are kept deliberately, because cross-confederation matches are the only games that let the model compare the relative strength of teams from different continents. The home advantage factor is removed for matches played at neutral venues.

Football Tournaments — Group Stage plus Knockout Bracket

Tournament competitions such as the World Cup and European Championship are not simple league tables, so their outright predictions (who reaches the final, who wins it) use a dedicated two-stage simulator built on top of the Dixon-Coles match engine described above. Each run simulates the full path to the trophy:

Aggregating thousands of these full-tournament runs produces, for every team, the probability of advancing from the group, reaching the quarter-finals, semi-finals, and final, and winning the tournament outright. These are the columns shown on a tournament's outright page. Because each run is considerably heavier than a single league simulation, tournaments use 5,000 runs per update rather than 10,000.

Formula 1 — Plackett-Luce Model

F1 predictions use a Plackett-Luce ranking model, which estimates each driver's relative strength and converts it into probabilities for every finishing position. Unlike simple regression, Plackett-Luce naturally handles the full finishing order and produces consistent multi-outcome probabilities (winner, podium, top 6).

The model considers: qualifying position (the strongest single predictor of race outcome), recent driver form across the last several races, constructor performance trends in the current season, driver history at each specific circuit, and session data from practice and qualifying.

Predictions are generated in two windows: a pre-weekend estimate using historical data only, and an updated post-qualifying prediction that incorporates the actual qualifying result and weekend session pace.

Data is sourced from the FastF1 Python library, which provides access to the official F1 Live Timing feed including qualifying times, race results, lap times, pit stops, and tyre compound information.

WRC — Surface-Adjusted Elo

WRC predictions use a surface-adjusted Elo rating system combined with Bradley-Terry probability modelling.

Each driver maintains separate Elo ratings for overall performance, gravel rallies, asphalt rallies, and snow rallies. Ratings are updated after every event using finishing order and opponent strength, with higher K-factors for newer drivers so the model adapts faster when limited historical data exists.

Rally win and podium probabilities are derived from the Elo ratings, while reliability is modelled separately using each driver's historical finish rate. Drivers with frequent retirements due to crashes or mechanical issues receive lower effective probabilities.

Key factors include surface-specific performance, recent form, historical results at each rally, and overall reliability.

Tennis — Elo Ratings for ATP Singles

Tennis predictions use a head-to-head Elo rating system. Every player carries a single overall rating; the probability that one beats the other is the standard logistic function of the difference between their ratings, on the conventional 400-point scale. A 100-point gap means roughly a 64% chance for the higher-rated player, a 200-point gap roughly 76%.

Ratings move after every completed match by an amount that depends on how surprising the result was and on how much the model already knows about each player. The K-factor — the size of that adjustment — starts at 32 for a newcomer and decays toward 20 over roughly the first 80 rated matches. A qualifier arriving from the Challenger circuit therefore moves quickly on his first results, while an established player's rating barely shifts on a single match.

Ratings are continuous across seasons. A player arrives in January carrying what he built over previous years rather than starting level with a debutant. Between seasons his match count is halved, which raises his K-factor again: the model treats an established player returning from a winter break as temporarily more uncertain, so his first few results of the year move his rating faster than they would mid-season.

Everything is rebuilt from scratch on every run rather than accumulated. The entire stored match history is replayed in chronological order each night, so a day the data collection missed leaves no permanent distortion — the backfilled match simply takes its place in the sequence.

Walkovers and defaults are excluded from both the ratings and the accuracy scoring: they are administrative outcomes, not sporting ones, and rating them would credit a player for an opponent's illness. Retirements are treated as genuine results and count normally.

Two markets are published: the winner of each individual match, and the winner of the tournament outright. Doubles and qualifying matches are excluded throughout.

Tennis — Tournament Outright

Once a draw is published, the full bracket is simulated 10,000 times. Every tie is decided by the same rating comparison used for individual matches, winners advance through the bracket exactly as they would in the tournament, and the proportion of runs each player wins becomes his probability of taking the title.

A bye advances a player without simulating anything. It is the absence of a match, not an easy one, and treating it as a match would invent a chance of losing that does not exist.

No outright is published until the draw is complete. Qualifying finishes the day before the main draw, so a bracket published earlier still has empty places. Filling them with an invented rating would not simply add noise: the first published figure is the one the model is scored against, so a guess there would become the benchmark for everything that follows. Waiting costs about a day out of a market that runs for one to two weeks.

The estimate is republished daily as the tournament progresses, and each update starts from the players still in the draw rather than re-running the original field. A player knocked out in the first round therefore disappears from the figures rather than lingering at a few percent all week. Only the first snapshot — the one made before a ball was struck — counts towards the accuracy measured on the Performance page; the later ones show how the picture changed, and scoring them would flatter the model, since a forecast made once two players are left is barely a forecast at all.

For the same reason, a tournament that Proknoz only starts following midway through gets no outright at all. There is no honest baseline to publish.

Tennis — Surface ratings, and why they are shown but not used

Each player also carries separate ratings for hard, clay and grass courts, updated alongside the overall rating and displayed on player and match pages. Indoor hard courts are treated as hard courts.

These surface ratings do not currently influence any prediction, and that is a deliberate decision based on measurement rather than an oversight. Two separate studies, on two different datasets, found that blending the surface rating into the prediction produced no improvement at all.

The second study also found the reason. A surface rating only begins to exist when a player first plays on that surface, and it then accumulates roughly a third as many matches as his overall rating — so it permanently lags behind. In the sample measured, around 74% of hard-court ratings and half of clay ratings sat below the player's own overall rating. Because a weaker player's rating already sits near the average, that lag pulled the stronger player down further than the weaker one, and the net effect was to compress the gap between them rather than adjust for surface. On clay it was destroying about a fifth of the signal while appearing to refine it.

Turning the blend off removed a systematic under-confidence in the published probabilities: before the change, the favourite won more often than the model said in every band of the rating scale. The surface ratings are still maintained and still shown, because they are informative to a reader and because a corrected version of the blend may earn its place in future.

Tennis — What the model does not know

A new player starts at the tour average of 1500 regardless of his ATP ranking. This was measured before being accepted: matches involving players with fewer than three rated results are about 9% of the sample, and even a perfect starting estimate for them could improve overall accuracy by roughly 0.001 — smaller than the measurement noise. The fix would cost data-collection budget that is better spent elsewhere.

The model also has no knowledge of the tournament's importance beyond the players involved. Weighting ATP 250 results as less informative than Grand Slam results was tested across several schemes and produced no reliable improvement. Head-to-head history, recent form and opponent quality are shown on match pages as context, but none of them feeds the prediction — the rating already reflects who a player has beaten.

Predictions cover ATP singles only. Doubles, WTA, Challenger and ITF events are not modelled, although Challenger results would be a natural way to give newcomers a realistic starting rating.

Cycling — Stage-Type Regression

Cycling predictions use a logistic regression model trained separately for each stage type: flat (sprint), mountain, time trial, and mixed/hilly. Treating stage types as distinct prediction problems reflects the reality that specialists consistently outperform generalists in their preferred terrain.

Each rider is represented by a feature vector combining specialty scores (derived from ProCyclingStats profile ratings), recent form, and race fatigue. The specialty scores measure a rider's ability in six dimensions: sprinting, climbing, time trialling, punching power, one-day racing, and GC ability. Stage-type win rate — the fraction of top-5 finishes in stages of the same type over the previous two years — is the most predictive single feature and ensures that sprinters dominate sprint stages and climbers dominate mountain stages regardless of overall prestige or UCI points.

Fatigue is modelled in two components. Intra-race fatigue accumulates kilometres and elevation within the current race and is normalised against a full Grand Tour (approximately 3,500 km and 50,000 m). Inter-race fatigue measures total kilometres raced across all events in the previous 28 days, capturing the accumulated load from a packed calendar. Both components influence both predicted performance and DNF probability.

At prediction time, raw model scores are converted to a proper probability distribution via softmax across the entire field, so probabilities sum to exactly 100% across all starters.

Top-10 probabilities are drawn from that same distribution rather than derived from it by formula. Ten thousand finishing orders are sampled, each one picking a winner in proportion to the rider weights, removing him, picking the next from the remainder, and so on; a rider's top-10 probability is the fraction of those runs in which he finished in the first ten. Sampling rather than approximating keeps the two markets consistent with each other, and keeps the published figures away from 0% and 100% — ten thousand runs that all agree are evidence, not certainty.

Probabilities are computed across the full startlist but only the thirty strongest candidates are published for each stage. The figures shown on a stage page therefore do not sum to 100%: the remainder belongs to the rest of the peloton, whose individual chances are too small to be worth listing.

Models are trained on four seasons of WorldTour results (2022–2025) covering Grand Tours, monuments, and major stage races. After each stage completes, the relevant stage-type model is partially retrained with the new result, keeping predictions fresh as form evolves through the season.

Data is sourced from ProCyclingStats via FlareSolverr (Cloudflare bypass), with a rotating proxy pool to ensure reliable access. Historical stage results, rider profiles, and specialty scores are cached locally to minimise scraping frequency.

Cycling — GC and Classification Predictions

General classification (GC) predictions for stage races use Monte Carlo simulation over the remaining stages. Starting from the current live GC standings, each simulation run samples a time gain or loss for every active rider on each remaining stage, using distributions calibrated by stage type and rider specialty. GC contenders (high GC score or within five minutes of the leader) receive tighter distributions; domestiques and non-GC riders receive wider distributions reflecting their inconsistency in the mountains.

DNF probability increases with accumulated fatigue and each rider's historical retirement rate, and is applied independently on every simulated stage. Riders who abandon in a simulation run are excluded from subsequent stages in that run.

Mountains and Points classification probabilities are approximated from the GC simulation results using specialty score weighting: climber score for KOM, sprint score for Points. White jersey (best young rider) probabilities are derived by renormalising GC win probabilities over under-23 riders only.

For one-day classics, no GC simulation is needed — the stage winner model is applied directly to the full startlist.

Outright Predictions — Monte Carlo Simulation

League winner and championship predictions use Monte Carlo simulation. Remaining fixtures or events are simulated 10,000+ times using the current model parameters, producing a distribution of final standings.

In WRC, each simulated rally includes overall classification, Super Sunday, and Power Stage points. Manufacturer simulations also apply FIA scoring restrictions such as limiting points to nominated factory entries.

These probabilities are updated after every matchday, race weekend, rally, or stage so users can track how predictions evolve throughout the season.

Reasoning Explanations

Each prediction includes an automated reasoning summary. This is generated from templates — the model identifies the 3–5 most influential factors for that specific prediction, maps them to human-readable sentences, and composes a brief narrative. No large language models are involved in this step.

Depending on the sport, explanations may reference recent form, surface-specific performance, injuries and suspensions, qualifying pace, reliability, stage type suitability, accumulated fatigue, expected goals quality, confirmed lineup formations, or historical performance at a venue or rally.

For tennis, the explanation describes the size of the rating gap between the two players, and may reference their head-to-head record, recent results, and form on the match surface. The wording of the gap is calibrated against what actually happens: across the matches measured, a gap the model calls decisive was won by the favourite about 81% of the time, one called clear about 69%, and one called slight about 59%. Below that the match is described as close, and the favourite won about 56%. Head-to-head and surface form appear in the explanation as context for the reader — they are not inputs to the probability.

Tournament outright explanations are deliberately narrower. The bracket simulation reads two things and no more: each player's rating, and the shape of the draw. So the explanation reports where a player stands on rating among the field, and how often the simulation takes him to the final and lets him win it — and says nothing about his seeding, his recent form or his head-to-head record. Those are true facts about a player, but none of them was consulted, and presenting an unread number as a reason for the published one would be misleading however accurate the sentence.

Calibration and Accuracy

The Performance page tracks every published prediction against what actually happened. Three things about how it counts are worth stating plainly, because each of them changes the numbers.

The unit is a decision, not a line of probability. Predicting the winner of a Tour de France stage means assigning a probability to every rider in the peloton, and almost all of those lines are easy calls that the rider in question will not win. Counting each one as a separate prediction would let a single stage outweigh fifty football matches. One stage, one match, one title race counts once, and the score for it is the average across its lines.

Raw scores are not comparable between markets. A Brier score depends on how hard the question is. Guessing blindly scores about 0.25 on a tennis match, where there are two outcomes, and about 0.006 on a stage with 160 starters, where almost every line is correctly near zero. A low number in a large field is therefore not evidence of a good model. The Performance page shows a skill score alongside the raw Brier: how much better the model does than a forecaster who knows only how often this kind of call comes in. Zero means it adds nothing, one is perfect, and a negative number means it is worse than that baseline. The baseline itself is measured from what has happened, not assumed, and is computed separately for each market and each field size — winning a six-rider classic and winning a Grand Tour stage are the same market and not the same task.

Some predictions are deliberately not counted, and in every case it is because scoring them would flatter the model rather than measure it:

Because the baseline is measured rather than assumed, the page can and does report that a model is worse than guessing. That is the intended behaviour. A tracking page that can only produce flattering numbers is not tracking anything.

Data Sources

All data comes from free, publicly available sources.

For football: Football-Data.co.uk (historical CSV data for model training), Football-Data.org (fixtures and results API), the martj42 international results dataset (national-team match history for World Cup and Euro modelling), Understat (per-team season xG averages for the top five European leagues), FBref (xG for Liga Portugal and European cup competitions), and Transfermarkt (player injuries, suspensions, and confirmed lineups — scraped from public league pages).

For Formula 1: the FastF1 library, which provides qualifying times, race results, lap data, pit stops, and tyre information from the official F1 Live Timing feed.

For WRC: the official WRC.com public API, which provides structured event data including itineraries, entry lists, stage results, rally classifications, and championship standings.

For cycling: ProCyclingStats, which provides race calendars, stage profiles, startlists, results, and rider specialty profiles for all WorldTour events.

For tennis: the tennis-api (Matchstat) service, which provides the ATP calendar, tournament results, match details, and the draw bracket for each tournament. Head-to-head records are not requested from the source — they are derived from the matches already stored, which covers every meeting Proknoz has recorded.

Limitations

Models are simplifications of reality. They don't fully account for tactical changes, transfer-window signings, managerial motivation, weather conditions beyond basic factors, or unpredictable events.

Football injury data from Transfermarkt reflects publicly reported absences, which may lag behind actual squad news by a day or two. Players listed as doubtful but who ultimately play are not automatically corrected before kick-off unless a confirmed lineup is available. The xG adjustment applies season-long averages and does not capture recent shooting form changes within the season.

For international football, comparing the strength of teams from different confederations is inherently difficult, because they rarely play each other outside major tournaments. The model leans on friendlies and intercontinental fixtures to calibrate this, and applies regularisation to limit over-rating, but some teams that dominate weaker regional opposition may still be rated more highly than a neutral observer would expect. Tournament outright probabilities should be read in that light, particularly in the early rounds when few matches have been played. As more matches are played and tracked on the Performance page, these estimates can be assessed against actual outcomes.

Tournament outright simulations currently treat all knockout matches as played at neutral venues and do not model travel, altitude, host-nation advantage, or squad rotation between fixtures. Extra time and penalty shoot-outs are modelled as a single tie-break rather than as separately simulated phases.

WRC predictions do not currently model road position effects on gravel, changing weather conditions during rallies, or team orders.

Tennis predictions know nothing about injuries, fitness, travel, altitude, court speed, weather, or a player's motivation for a particular event — all of which matter in a sport decided by one athlete against another. A rating summarises results and nothing else. Ratings for players new to the tour are unreliable until they have accumulated a season or so of results.

The outright simulation inherits every one of those blind spots and compounds them, because a tournament is six or seven matches and the errors accumulate along the path. It also assumes the draw is played as published: a withdrawal after the bracket is set is not reflected until the following day's update. And it does not distinguish a favourable draw from an unfavourable one in the explanation, because measuring that properly requires comparing the actual bracket against an average one, which is not yet done.

Cycling predictions do not model team tactics (leadout trains, domestique support, protected riders), breakaway probability, or intermediate sprint bonuses in the GC simulation. One-day race predictions treat each event independently and do not account for fatigue from earlier races in the same week (e.g. the Ardennes classics).

Cycling stage predictions rank a field the model has seen only through results and profile ratings. A rider in unusual form, a team riding for a team-mate, or a breakaway that is allowed to stay away will all beat the model, and the winner of a stage is frequently outside the thirty riders it considered most likely. The Performance page measures how often that happens rather than hiding it.

Probability estimates are the model's best assessment, not certainty.

The models are trained primarily on recent seasons and may perform less well for newly promoted teams, rookie riders, or competitions with limited historical data.