Pose tracking can estimate what players see before receiving

Pose tracking can estimate what players see before receiving

Scanning is no longer just a head-count metric. A pose-tracking model estimates what space a player was likely able to see before the ball arrived, then tests whether those visual-access features predict clear changes in his imminent pitch value.

The next step beyond counting scans

Most scan analysis starts with a simple event: did the player move his head before receiving? In the work summarized by The xG Football Club, Bekkers operationalizes traditional Visual Exploratory Action features as rapid head movements where head angular velocity exceeds 125°/s.

That has coaching value. The next step is adding context. A scan count does not tell you whether the player saw the spare man, the pressing defender, the covering centre-back, or only empty grass. It also turns perception into a binary event: scanning or not scanning.

Joris Bekkers’ paper, “Wide Open Gazes”, proposes a different unit of analysis: not “how many times did he look?” but “which areas of the pitch were likely visible to him while the ball was travelling?”

How the model builds a vision map

The model creates a continuous “stochastic vision layer” from pose-enhanced tracking data. It uses head and shoulder orientation to estimate where a player is looking, then projects that into a two-dimensional pitch map.

The field of view starts from a 120° binocular vision area, then becomes probabilistic. Vision is strongest centrally, weaker in the periphery, weaker at distance, and affected by player speed. The model also leans on the assumption that head orientation approximates gaze direction, supported by general eye-head coordination evidence that about 92% of gaze shifts are in the same direction as head movement; that is not the same as soccer-specific gaze validation.

The second layer is occlusion. Other players block sightlines. Each teammate or opponent creates a probabilistic shadow whose size depends on distance and shoulder orientation. A player facing the observer presents a larger obstruction than one angled sideways.

The result is a 105×68 pitch grid where every location has a probability of being visible to the receiving player. The accompanying open-source repository describes the same two-part construction: a field-of-view model plus an occlusion model, combined with pitch control and pitch value.

That is the important step. The model is not just counting body movement. It is asking whether the player likely had visual access to space that mattered.

The receiving phase is the test

The study focuses on the “awaiting phase”: the period after a teammate passes and before the receiver takes possession. That is a clean, measurable version of the coach’s “check your shoulder” window: the final receiving-specific update before first touch.

The dataset covers 32 Copa América 2024 matches, using broadcast tracking, pose estimation and synchronized event data. It includes about 2.3 million frames and more than 60,000 events, filtered to open-play sequences as defined by the paper: sequences with at least one successful pass and actions at least seven seconds after the most recent set piece. From that, roughly 14,000 awaiting moments were extracted.

Bekkers then links the vision map to two familiar football analytics surfaces. Pitch control estimates which team or player can influence each part of the pitch, while pitch value estimates how strategically valuable each location is given the ball position. The paper also introduces “imminent pitch control,” reducing the influence radius so the model focuses on space a player can affect in the very short term, not theoretical space over a longer horizon.

For validation, XGBoost classifiers predicted whether the receiver’s imminent pitch value clearly increased or decreased by the end of the subsequent on-ball phase, with ambiguous middle cases omitted. The comparison is the key result: the regular model reached an AUC of 0.744, while the vision-enhanced model reached 0.788. Adding traditional VEA scan features did not improve performance in that setup.

The useful signal was not “seeing more”

The most interesting finding is not that players should look around more. It is that the model’s useful features related to the quality of observed space.

The feature-importance results point less toward seeing more grass and more toward seeing relevant occupied space. Positive predictors included observed defensive occupied space, the defensive-to-attacking occupied-space ratio, observed attacking occupied space, and observed attack-controlled space. By contrast, observing more defensive-controlled space was associated with worse pitch-value outcomes, possibly because it marks highly contested situations.

That fits football logic. A midfielder receiving between lines does not need a panoramic tourist photo. He needs to know where the nearest pressure is, whether the cover shadow blocks the forward pass, whether the far-side defender has jumped, and whether the third-man option is open.

This moves scanning analysis from effort to information. The question becomes not “did he check?” but “did his check include the constraint that shaped the next action?”

Earlier scanning research still matters

This does not make scan-count research useless. It shows where the next layer of analysis can sit on top of it.

In a 2018 11v11 study, McGuckian and colleagues used head-mounted IMUs with 32 semi-elite players and analysed 783 possessions. Head-turn frequency and excursion rose as players got closer to receiving: mean frequency was 1.44 turns per second in the final second before possession, compared with 0.95 over the 10-second window. Higher exploration was associated with turning with the ball, playing forward passes, and playing passes to the opposite side from where the ball was received. It was not associated with successful pass completion.

The EPL PlayerCam study presented at the MIT Sloan Sports Analytics Conference also found a positive relationship between visual exploratory behaviour before receiving and on-ball performance, using 1,279 situations from 118 Premier League midfielders and forwards; the reported effect was largest for midfielders playing forward passes.

Context can make the metric more useful. A UEFA U17 and U21 field study on exploratory behaviour and passing found that the direction of the last scan before receiving related to the foot used for first contact and the direction play continued, with pressure and dominant foot also affecting the relationship.

That is exactly why a continuous vision map is attractive. Frequency, direction, body orientation, pressure, space value and opponent location are not separate coaching questions. They are one receiving situation.

Why this should not become a scouting shortcut

The useful caveat is that scan frequency alone should not be treated as proof of superior football intelligence.

A study published online in 2024 comparing “super elite” award-winning Champions League players with elite teammates found no significant differences in VEA frequency or performance when controlling partly for match dynamics by matching players from the same team, match and positioning line. The players scanned more during the final pass than the penultimate pass, but VEA did not distinguish the two groups.

That does not contradict Bekkers’ work. It makes the pose-tracking approach more interesting. If simple scan frequency does not reliably separate elite from super elite, the better question is what was likely visible, when it was likely visible, and under what pressure.

Why the estimate matters

Pose tracking does not show the exact contents of a player’s mind. It estimates likely visual access, which is still a meaningful upgrade on counting whether the head moved.

That matters because player orientation estimation is hard. These studies are caveats about the general difficulty of estimating orientation from football video, not direct validations or invalidations of Bekkers’ exact data pipeline. In a 2020 computer-vision paper on estimating footballer orientation from monocular video, researchers combined OpenPose-derived shoulders and hips with ball context, then validated against player-held EPTS devices, reporting a median error of 27 degrees per player. A 2025 body-orientation pipeline tested on Women’s Super League clips reported 75% accuracy, with errors linked to player, ball and pass-detection issues.

That is why the strongest wording is “probable visual access,” not a perfect replay of gaze. Head angle is not eye gaze. Occlusion models are approximations. Broadcast tracking and pose estimation introduce errors. The value is that the model connects estimated visual access to pitch-value outcomes without pretending to read the player’s mind.

What coaches should do with it

The practical value is in sharper review language.

Instead of praising a midfielder because he scanned three times, an analyst can ask whether the look happened during the pass travel or too early, whether the body orientation kept both the ball and the next pressure in view, whether the scan included the defender who could jump, and whether the information gathered matched the action chosen.

That is a better coaching loop. It connects the check before receiving to the first touch, the turn, the pass angle and the space gained or lost after the action.

The conclusion is narrow and useful: pose-enhanced tracking can make scanning analysis more specific. The future metric is not only “heads turned per second.” It is whether the player had access to the part of the game that mattered before the ball arrived.

2026

August

July

June

May

March

January

2025

December

November

October

September

August

July

June

May

April

March

February

January

2024

December

November

October

September

August

July

June

May

April

March

February

January

2023

December

November

October

September

August

July

June

May

April

March

February

January

Receive every new post in your inbox.