A transparent view of how LiveWin turns match evidence into probabilities, how confidence differs from certainty, and how every model release should earn its place.
Current system
Model version
livewin-hybrid-v0.4
Fixtures evaluated
15
League coverage
15
Markets observed
4
Odds snapshot
Jul 29, 2026, 4:18 AM
Provider status identifies the configured feed, not an independent audit of data completeness or model performance.
Forecast architecture
Recent form, season output, xG, shots, clean sheets, venue split, head-to-head and timestamped odds.
Attack and defensive rates produce matchup-specific scoring intensities and a scoreline probability matrix.
Statistical output is pooled with de-vigged market priors in log-odds space at a weight fitted on a 7,691-match backtest. The market carries most of it — measured, not assumed.
Confidence without theatre
A 58% home-win probability describes the forecasted event. Confidence describes how much the model trusts that estimate given data volume, signal agreement, matchup volatility and market divergence.
Forecast buckets should settle near the diagonal. A strong model is honest about uncertainty, not merely right often.
Release standard
| Metric | What it measures | Promotion rule |
|---|---|---|
| Log loss | Punishes confident mistakes and rewards complete probability distributions. | Must beat the previous model on an untouched time split. |
| Brier score | Measures probability error across home, draw and away outcomes. | Must improve overall and avoid material league-level regression. |
| Calibration | Checks whether events forecast at 60% occur close to 60% over time. | Reliability curve remains inside the defined tolerance bands. |
| Closing-line value | Tests whether identified prices tend to beat the market close. | Reported separately from hit rate and never presented as guaranteed return. |
Every result is a probability forecast. Confidence and value indicators are evidence summaries, not promises.
Generated analysis is constrained to structured inputs and must not fabricate injuries, lineups or table facts.
Development metrics and historical simulations remain clearly separated from independently audited live results.
Confidence measures evidence agreement; value requires a meaningful edge after market margin is removed.
Primary score
Log loss
Probability check
Brier score
Validation
Time-split