Validate football models by the misses, not the average

Validate football models by the misses, not the average

Pass detection reached 0.92 F-score on dependent data and 0.71 on independent data. Staff should see per-class misses, timing errors, and context splits before using any model score in a report.

Bischofberger, Baca and Schikuta tested pass and shot detection on four football datasets. 1 The independent Euro sample gave the stress test because its positional feed came from Tracab and its event feed came from Wyscout after manual timestamp correction from broadcast footage.

Pass F-scores reached 0.87 to 0.92 on datasets where event and positional data appeared linked. On Euro, where the paper says event and positional data were independent, passes fell to 0.71 and shots to 0.65.

Ball coordinates can already carry manual event cues. If the ball sticks to the possessing player in a provider feed, a detector may match that provider’s tagging habits instead of finding passes or shots on a new feed.

The dependent data drop

Dataset Sample Event and positional link Staff read
Metrica 3 games Linked by inspection High pass scores may include provider artifacts
Stats 14 games Linked to a lesser degree Good scores need provider-context checks
Euro 4 games Independent feeds Best stress test in the paper
Subsequent 6 games Linked to a lesser degree Parameters may not transfer cleanly
Event-detection F-scores by dataset showing independent data exposes the real error rate
Figure 1.1 - Independent data exposes the real error rate for pass and shot detection

On Euro, the authors estimated that roughly one third of detected passes or shots would not appear in the manual event data analysts are used to, and roughly one third of manual events would be missing.

Passes are about 40 times more common than shots, so a model can look good on the common action while failing on the rarer event that decides a clip reel, xG input, or chance-creation report.

What staff should ask for

In a dataset where positive examples appear 1 percent of the time, a model that predicts negative every time scores 99 percent accuracy and still detects none of the target cases. 2

Precision answers how many flagged examples are real. Recall answers how many real examples were found. F1 penalizes a model when one of those two collapses, which is why it is a better first check than accuracy for many rare football actions.

Injury datasets usually contain many more non-injury than injury examples. Averaged scores can miss the class that matters, so injury-class recall and F1 are the first check, alongside precision and specificity. 3

Staff request Football question it answers
Base rate for each class How rare is the action or outcome
Per-class precision How many flagged clips or risk alerts are real
Per-class recall How many real events the model missed
Per-class F1 Whether precision or recall is collapsing
False-positive examples Which wrong clips, players, or alerts staff would see
False-negative examples Which real events staff would never review
Threshold used Which error trade-off the analyst selected
Data-context split Whether the score holds in the feed staff will use

A false positive shot creates a clip the opponent did not actually create. A false negative cutback run removes evidence from a recruitment filter. A false injury flag can change load or availability discussions before anyone has checked the player.

Event classes and timing windows

SoccerNet’s action-spotting benchmark asks systems to locate 17 football classes, including penalties, goals, offsides, shots on target, shots off target, fouls, corners, substitutions and cards. Its tighter evaluation uses one-to-five-second temporal tolerance rather than five-to-60 seconds. 4

A detector that finds kick-offs and corners accurately can still miss fouls, offsides, red cards, or shots off target.

The PLOS study matched constituent events within 500 milliseconds for Stats, Metrica and Subsequent, and within 1000 milliseconds for Euro.

If a model places a shot tag after the block, the video search may miss the striker’s body shape and pressure before contact. If a foul tag lands too early, the analyst may review the approach and miss the actual contact.

Tracking needs context splits

FIFA’s EPTS programme says public test reports describe positioning and velocity accuracy in different velocity brackets. FIFA’s live player and ball tracking validation also notes 2D ball-position accuracy, live latency and future velocity and 3D reporting. 5 6

Stadium changes, extreme weather, special lighting and poor image quality can sit outside the training data and break tracking-derived outputs. 7

Output Split to request Why staff need it
Player position Camera feed, stadium, image quality Shape and spacing models depend on coordinates
Velocity and sprint data Velocity bracket High-speed errors matter more for physical reports
Ball position Ball speed, 2D or 3D, latency Pressing and possession models depend on live ball accuracy
Broadcast tracking Weather, lighting, occlusion Edge cases can create bad coordinates before analysis starts

Sprint-load reports need the high-velocity error bracket. Pressing models need the ball-latency figure. Rest-defence analysis needs player positions in crowded, partially occluded frames.

Injury flags show the cost of each miss

Freitas and colleagues studied 34 male professional players at a Portuguese first-division club. 8 Injury events were only 0.20 percent of observations.

The study reported 74.22 percent overall accuracy for its best short-time-frame approach, but the usable figures were class-level. The cost-sensitive SVM reached 71.43 percent sensitivity and 74.19 percent specificity, which tells staff both how often injury events were found and how often non-injury cases were kept clear.

A false positive can sideline a player unnecessarily or pull medical resources toward the wrong case. A false negative can leave a player in full training when the model was supposed to flag short-term risk.

The football handoff

Event data goes through A, B and C audits — automated detection, then manual review, then expert correction — before reaching recruitment and performance workflows. 9

Contextual data feeds set-play review, open-play trends, coach communication and recruitment filtering. 10 A missed line-breaking pass changes the opponent report. A poor cutback-receiver label can remove a centre-forward from a shortlist.

Staff use Minimum error view before use
Opposition clips Per-class precision and recall for the searched event
Set-play review Timing-window errors and false positives by restart type
Recruitment filter False negatives for the movement or action being filtered
Tracking report Position and velocity errors split by match context
Injury-risk alert Sensitivity, specificity and threshold choice

Public evidence is strongest for passes, shots, action spotting, tracking validation and football injury classification. It does not set one universal minimum F1 for every model a club might use.

Start with the football class, then the context split, then the error examples. The headline score becomes useful after those three are in front of you.

Open the last false-negative shots, false-positive corners, late foul tags, noisy high-speed runs and missed injury flags. If the misses come from the same competition, camera feed, velocity bracket and time window staff will use next week, the score has football meaning.

Sources

  1. PLOS One: Detecting football events in heterogeneous data sources
  2. Google: Classification: Accuracy, recall, precision, and related metrics
  3. Sports Medicine Open: Machine learning in football injury prediction — a systematic review
  4. SoccerNet: Action spotting
  5. FIFA: Electronic performance and tracking systems
  6. FIFA: Validation of ball tracking technology and real-time tracking data extends
  7. Hudl StatsBomb: Creating better data — AI homography estimation
  8. PLOS One: Predicting injury in professional football using machine learning and GPS data
  9. Hudl StatsBomb: StatsBomb data FAQ
  10. Hudl StatsBomb: Using StatsBomb 360 data as a performance analyst

2026

August

July

June

May

March

January

2025

December

November

October

September

August

July

June

May

April

March

February

January

2024

December

November

October

September

August

July

June

May

April

March

February

January

2023

December

November

October

September

August

July

June

May

April

March

February

January

Receive every new post in your inbox.