Case study · Computer vision

Reading the game: live win probability from broadcast video alone.

Independent R&D, February 2025. No API, no game data feed: just the pixels a viewer sees. The pipeline detects players on the minimap, reads the score and clock from the HUD, and predicts the winner every second of the match.

Source footage is publicly broadcast Call of Duty League match video used for research. Team names and logos visible in screenshots belong to their owners.

99.3%

frame coverage after post-processing

28% → 97%

hold-out accuracy, raw vs. cleaned data

111,360

frames processed

~90%

labeling time saved with model-assisted annotation

Demo clip

Fifteen-second excerpt of the rendered overlay. The bar at the bottom is the model's live estimate; the score and clock it reads are in the HUD at the top.

The problem

Real-time win probability with no data feed

Traditional sports broadcasts show a win-probability graphic driven by official play-by-play data. Esports broadcasts have no equivalent feed available to a third party. Everything the model needs, the score, the clock, and who controls the map, is only visible as pixels on screen.

The goal was a system that ingests standard 1080p/60fps broadcast video of Call of Duty Hardpoint matches and outputs a calibrated probability, every second, of each team winning. It had to be reliable enough that the overlay never shows an impossible state.

Minimap crop with YOLOv8 bounding boxes: four green boxes labeled player with confidences 0.74 to 0.87 and one red box labeled enemy with confidence 0.43
YOLOv8 detections on the minimap: friendly markers in green, an enemy in red.

The pipeline

Four stages, video in, probability out

Pipeline diagram: broadcast video input feeds minimap detection (YOLOv8 player counts) and scoreboard OCR (template matching); both feed feature engineering with 12 features, then an XGBoost classifier, then a broadcast overlay
Architecture diagram of the pipeline.
  1. Data acquisition and labeling

    Frames sampled every five seconds from 1080p/60fps match video, cropped to the fixed minimap region. Fifty frames were labeled by hand in LabelMe; a small YOLOv8-Nano model trained on those fifty auto-labeled the remaining ~1,450 for human correction, cutting annotation time by roughly 90%.

  2. Minimap detection (YOLOv8-Medium)

    A YOLOv8-Medium detector trained for 100 epochs on Apple Silicon (MPS) finds friendly and enemy markers on the minimap. Counting detections above a 0.25 confidence threshold gives map-control telemetry the scoreboard cannot provide.

  3. Scoreboard and clock OCR (template matching)

    Scores (0–250) and the game clock are read from fixed HUD regions with OpenCV normalized cross-correlation against digit templates extracted from the video itself. Instead of reading arbitrary digit strings, the system enumerates every valid score and picks the best-fitting one, which rules out impossible readings by construction.

  4. Win-probability model (XGBoost)

    Twelve engineered features (six scoreboard, six minimap, including 30-second rolling map-control averages) feed an XGBoost classifier. Mirror augmentation swaps team perspectives to double the training rows and balance classes. Output is smoothed with a 45-second rolling average and rendered as a broadcast overlay at 59.94 fps.

LabelMe annotation tool showing a minimap frame with hand-drawn player and enemy bounding boxes, frame 7 of 732
Seed labeling in LabelMe. Fifty hand-labeled frames were enough to bootstrap model-assisted annotation of the rest.

What was hard

Where the time actually went

Neural OCR hallucinated plausible-looking numbers

The first pass used EasyOCR. It read "250" as "2S0", produced scores of 839 in a game capped at 250, and let scores go backwards. The errors looked fine at a glance and quietly corrupted the training data. Replacing it with template matching plus score enumeration removed impossible readings entirely.

Anti-aliasing made identical digits match differently

Sub-pixel anti-aliasing varies frame to frame, so a template cropped from one frame scored poorly on the same digit elsewhere. The fix was mundane: up to ten template variants per digit, extracted at native resolution from frames with a known score, and matching on the whole region rather than segmenting tightly kerned digits first.

The model was fine; the data was not

Raw extraction read 71% of score frames and 42% of clock frames. Outlier removal, monotonicity enforcement (Hardpoint scores only go up) and forward-fill brought coverage to 99.3%. A row-level audit then threw out 18 of 21 candidate match files for impossible score jumps or a 250–250 final. Hold-out accuracy went from 28.2% on the raw data to 97.2% on the cleaned data with the same model.

Results

99.3% coverage, 97% hold-out accuracy

Validation used a hold-out match the model had never seen. The top panel is the actual score of both teams; the bottom panel is the model's probability that Team A wins, computed only from what it read off the screen. The probability crosses 50% about forty seconds before Team A takes the lead on the scoreboard, because the minimap features see map control shift before it turns into points.

Two-panel chart for the hold-out match. Top: Team A (red) and Team B (blue) scores rising over 575 seconds to 250 and 166. Bottom: Team A win probability near 0% for the first 210 seconds, climbing past 50% around 255 seconds and holding near 100% from 410 seconds onward
Hold-out match: actual scores versus predicted win probability.
Early-iteration win-probability chart over a 3,800-second match with the probability swinging sharply between 0% and 100% many times, final score 213 to 3
Before the data audit: an early model trained on uncleaned OCR output. The swings are artifacts of bad reads, and the 213–3 final is not a real Hardpoint result.

Top features by importance

  • 0.26 score_diff
  • 0.16 player_alive_rolling
  • 0.15 enemy_alive_rolling
  • 0.13 advantage_rolling

Roughly 43% of predictive power came from minimap telemetry, information that does not exist anywhere in the scoreboard.

Known limits

Three clean matches in the training set; some residual 5/6 and 0/8 digit confusion masked by smoothing; templates and regions calibrated for one broadcast layout. All documented in the validation report.

What transfers to agency work

The same methods, less glamorous video

A game broadcast is a convenient stand-in for any camera pointed at something that reports its state visually. The pieces that mattered, fixed-region reading, small-object counting, temporal smoothing and a validation gate, are the pieces agency problems need.

Gauge and panel reading where no data bus exists

Legacy analog gauges, seven-segment displays, control panels and HUDs can be read from a camera the same way the scoreboard was: fixed regions, template matching for stylized fonts, and validation rules that reject physically impossible values.

Counting from fixed or drone video

The minimap detector is a small-object counting problem. The same approach applies to vehicles in a lot, people at an entrance, animals in a survey transect or equipment in a yard, with rolling averages to smooth frame-level noise.

Telemetry extraction from recorded video

Archived inspection, range or training footage becomes a timestamped table of readings, counts and positions that analysts can query, chart and join to other records.

Document and form OCR with validation rules

The lesson that OCR output needs domain constraints (caps, monotonicity, allowed values) carries directly to digitizing forms and legacy records where a confident wrong read is worse than a blank.

Tech stack

Tools used

Python 3.12
Pipeline language
PyTorch + MPS
Training and inference on Apple Silicon
Ultralytics YOLOv8
Nano for labeling assistance, Medium for production
OpenCV
Video I/O, thresholding, template matching (TM_CCOEFF_NORMED)
XGBoost
Win-probability classifier
scikit-learn
Accuracy, AUC-ROC, hold-out evaluation
pandas
Telemetry cleaning and feature engineering
LabelMe
Bounding-box annotation
yt-dlp
Source video acquisition
matplotlib
Validation charts and overlay rendering

Have video you need numbers from?

Send a sample clip and a sentence about what you need to count or read. We will tell you honestly whether it is a two-week job or a research project.