Source-linked AI summary

A Comprehensive Review of Computer Vision in Sports: Open Issues, Future Trends and Research Directions

Banoth Thulasya Naik, Mohammad Farukh Hashmi, Neeraj Dhanraj Bokde

arXiv:2203.02281v2cs.CVeess.IV

TL;DR

Sports-video analysis must support detailed operations across diverse sports despite difficult visual conditions and limited quantitative benchmarking. This paper comprehensively reviews computer-vision applications, datasets, platforms, challenges, and research directions, reporting application-specific results including 97.56% real-time accuracy, 83.3% win prediction, and 72.7% loss prediction. It concludes that sports vision remains an open research area requiring improved data and evaluation.

  • Problem

    Sports vision involves challenging tasks such as player similarity, blurry video, sudden movements, and severe occlusion, while quantitative algorithm benchmarking is difficult because suitable benchmarks are unavailable.

  • Method

    The paper comprehensively reviews sports-video computer-vision applications, datasets, AI applications, GPU-based platforms, embedded platforms, challenges, and research directions.

  • Results

    Reported soccer results include 97.56% real-time accuracy, 83.3% accuracy for wins, and 72.7% for losses.

  • Takeaways & Limitations

    The review identifies sports vision as an open research area and provides research directions for addressing its challenges.

  • Takeaways & Limitations

    Quantitative benchmarking remains difficult because a suitable performance benchmark of algorithms is unavailable.

Abstract

from arXiv · show

Recent developments in video analysis of sports and computer vision techniques have achieved significant improvements to enable a variety of critical operations. To provide enhanced information, such as detailed complex analysis in sports like soccer, basketball, cricket, badminton, etc., studies have focused mainly on computer vision techniques employed to carry out different tasks. This paper presents a comprehensive review of sports video analysis for various applications high-level analysis such as detection and classification of players, tracking player or ball in sports and predicting the trajectories of player or ball, recognizing the teams strategies, classifying various events in sports. The paper further discusses published works in a variety of application-specific tasks related to sports and the present researchers views regarding them. Since there is a wide research scope in sports for deploying computer vision techniques in various sports, some of the publicly available datasets related to a particular sport have been provided. This work reviews a detailed discussion on some of the artificial intelligence(AI)applications in sports vision, GPU-based work stations, and embedded platforms. Finally, this review identifies the research directions, probable challenges, and future trends in the area of visual recognition in sports.

1.1 Features of the Proposed Review

This review synthesizes computer-vision sports research across detection, tracking, recognition, datasets, platforms, applications, and future directions. It extends prior sport-specific surveys with statistics, algorithm-selection guidance, evaluation criteria, datasets, and implementation considerations.

  • The paper also situates related work on badminton analysis, sports data mining, action recognition, video summarization, wearable technology, and inertial sensors.
  • Prior reviews addressed motion-capture systems, soccer player tracking, sports applications, movement recognition, ball tracking, content analysis, and sports AI.
  • The review covers sports-video tasks including player and ball detection, tracking, trajectory prediction, team-strategy recognition, and event classification.
  • Unlike recent reviews, it analyzes study statistics across sports and AI algorithms used for multiple sports-vision aspects.
  • It provides an AI-algorithm selection and evaluation roadmap, publicly available sports datasets, GPU-based embedded platforms, and research directions.

2 Statistics of Studies in Sports

The section characterizes recent sports-vision research and its applications, while linking player-position detection and movement analysis to real-time team insight. It surveys publication progress and sport-specific application statistics.

  • Player-position detection is the basic step for tracking a player.
  • Analyzing individual movements and real-time team formations can provide coaches with real-time insight into team performance.
  • Sports-vision applications include player, ball, and referee detection and tracking, object classification, performance analysis, gesture recognition, highlight detection, and score updating.
  • Figures 3 and 4 summarize sports-research publications over the past five years and studies across sports and applications.

3 Play Field Extraction in Various Sports

Playfield extraction supports downstream sports-video analysis but remains sensitive to lines, illumination, shadows, viewpoints, and appearance changes. The review compares color, background-subtraction, labeling, and camera-based approaches.

  • Playfield extraction separates the playing region from non-playfield areas and supports player or ball detection, tracking, event extraction, and pose detection.
  • Gaussian background subtraction generates foreground masks by detecting moving objects, but background-subtraction methods fail to detect playfield lines.
  • Color-based methods use RGB-to-HSI, YCbCr, or normalized-RGB transformations to reduce illumination effects.
  • Accurate capture requires camera calibration and sufficient camera coverage, while challenging views include occlusion, missing balls, shadows, weather, and changing illumination.
  • Labeling playfield lines, advertisements, and non-playfield regions enables their detection and classification while reducing false positives and false negatives.

4 Literature Review

The literature review surveys computer-vision methods for sports detection, recognition, tracking, trajectory prediction, and performance analysis. Across sports, reported systems address diverse tasks but remain constrained by dynamic play, occlusion, identity ambiguity, and difficult evaluation.

  • Basketball: Basketball studies address event classification, pose and location detection, defensive-strategy recognition, player-position prediction, and three-dimensional ball-trajectory generation.
  • Basketball: Basketball analysis is difficult because game situations change rapidly and player behavior is dynamic, complicating play-level tracking and trajectory prediction.
  • Basketball: Reported basketball results include 69% classification accuracy for automatic defensive behavior and a final AUC of 91%.
  • Common limitations include severe occlusion, same-colored jerseys, tracking ambiguities, identity switches, inconsistent performance under unusual views, and trajectory failures in uncertain scenarios.
  • Soccer: Soccer systems support player and ball detection, jersey-number identification, tracking, action recognition, pass classification, tactical analysis, and trajectory prediction.
  • Soccer: Soccer reports include 97.56% accuracy with real-time operation, 83.3% accuracy for wins and 72.7% for losses, and 87.6% and 88% F1-scores.

4.3 Cricket

Computer vision studies in cricket address batting, bowling, ball movement, umpire decisions, video summarization, and player-performance prediction. Reported methods achieve strong results across several tasks, while dataset availability, class imbalance, and coverage constraints remain recurring limitations.

  • Cricket video analysis covers batting shots, bowling performance, scoring, ball trajectories, umpire decisions, commentary generation, and player-performance prediction.
  • Video summarization selects interesting cricket events and produces compact representations of match footage, but summarized methodologies retain application-specific limitations.
  • 96.82%, 94.83%, and 96.32% were reported for precision, F1-score, and accuracy across four cricket classes using a decision-tree methodology.
  • Cricket systems remain constrained by duplicate or imbalanced data, unavailable standard ball-outcome datasets, frame-rate dependence, and incomplete coverage of bowler types.
  • 93.73% classification accuracy was achieved by replacing machine-learning techniques with deep-learning techniques in one cricket study.
  • 94% accuracy with DCNN and 100% with Inception V3 were reported for umpire-signal classification toward automated scoring.

4.4 Tennis

Tennis computer vision research spans tracking, shot and activity recognition, trajectory analysis, player behavior, and match-level visualization. Reported systems achieve useful classification and detection results, but performance depends on movement similarity, training coverage, computational resources, and match conditions.

  • Classification: 20 classes were identified by one proposed tennis method, while an average F1-score of 88% was reported for another classification task.
  • Applications: Tennis analysis uses ball and player tracking for match visualization, trajectory prediction, activity recognition, behavior analysis, and next-shot prediction.
  • Limitations: Tennis models can fail when movements are similar, player styles differ, frame rates change, or multi-object tracking is required.
  • Evaluation: Performance factors can be measured from the minimum distance between predicted and ground-truth shot locations.
  • Results: 84.39%, 75.81%, and 79.87% were reported for precision, recall, and F1-score on the Australian Open tennis match.
  • Results: 90.7% accuracy was achieved for tennis videos compared with 87.6% for badminton videos by the proposed algorithm.

4.7 Badminton

Badminton computer vision studies target player actions, performance, pose, shuttlecock tracking, and rally segmentation. The reviewed methods report strong results for selected tasks, while shuttlecock visibility, environmental conditions, overlap, scaling, and dataset limitations constrain robustness.

  • Applications: Badminton video analysis includes action recognition, player-performance analysis, player and shuttlecock tracking, and rally-scene segmentation.
  • Limitations: Badminton systems may fail under changing environmental conditions, binocular-camera constraints, shuttlecock detection difficulties, and player overlap.
  • Event detection: 95.8% average accuracy was reported for key-event and replay detection in sports video analysis.
  • Player tracking: 94% accuracy was reported for tracking overlapping players using an advanced optimization and matching approach.
  • Classification: 92% accuracy was achieved for classification across ten types of environmental conditions.
  • Deep learning: A 3 Dimensional CNN achieved 81% accuracy, while overlapping players reduced tracking performance because of similar features.

5 Available Datasets of Sports

The review catalogs publicly available sports datasets spanning videos and still images, with annotations supporting player, ball, referee, pose, event, and activity analysis. Shared datasets enable comparisons, common evaluation, and more transparent research.

  • Sharing annotated datasets supports comparison, common performance assessment, research transparency, and reduced effort in collecting and annotating large video quantities.
  • Public sports datasets include videos and still images captured with static or moving cameras for player-action recognition, event detection, and classification.
  • Their annotations cover player and ball tracking, pose estimation, sports-event summarization, referee movement, baskets, and other event parameters.
  • ISSIA annotates ball, player, and referee positions in each video from each camera, while TTNet labels ball bouncing, ball-net contact, and empty events.
  • APIDIS captures basketball videos from seven cameras positioned above and around the court, annotating player positions, referee movements, baskets, and ball position.

6 GPU-Based Work Stations and Embedded Platforms

The review compares GPU-based workstations, FPGA systems, and embedded platforms used for sports computer-vision tasks. It emphasizes the trade-offs among accuracy, computation speed, model size, and deployment constraints.

  • NVIDIA Jetson devices provide low-power, high-performance support for AI-based visual recognition and are configured with OpenCV, cuDNN, CUDA, L4T, and TensorRT.
  • A stereo-vision system tracked an indoor golf ball moving at 360 km/h using an FPGA board with an ARM CPU.
  • Real-time deployment must balance accuracy, computation speed, and model size across hardware platforms.
  • Sports vision implementations use GPU workstations, FPGA boards, GPU-constrained devices, and embedded platforms for detection, tracking, recognition, and motion analysis.

7 Applications in Sports Vision

Sports vision applications span fan engagement, coaching, officiating, tactical analysis, player tracking, and broadcast enhancement. The review describes AI systems as supporting richer analysis and increasingly automated decisions.

  • AI in the Sport Industry: AI supports fan-facing services, including personalized experiences, live game information, virtual assistants, and content that can increase engagement and create revenue opportunities.
  • AI in the Sport Industry: Sports applications analyze athlete performance and assist coaches with team guidance, opponent tactics, training, and dissemination of accumulated coaching knowledge.
  • Commercial Sports Technologies: Computer vision and AI are used for ball tracking, officiating, player tracking, pose estimation, tactical analysis, and enhanced sports broadcasting.
  • AI-Assisted Officiating: In tennis, computer vision can detect ball placement and speed, eliminating the need for a line umpire, while future systems may assist officials through earpieces and glasses.
  • Commercial Sports Technologies: Commercial systems such as SportsVu provide real-time optical tracking and comprehensive player data for tactical match analysis.

8 Research Directions in Sports Vision

The review identifies open problems in sports vision involving tracking, occlusion, pose, ball detection, trajectory prediction, jersey recognition, and real-time deployment. It proposes richer temporal data, identity-aware models, and task-specific datasets as research directions.

  • Research Directions: Future sports-vision research is organized around player and ball detection, tracking, pose estimation, trajectory prediction, and related task-specific applications.
  • Player Tracking: Deep representations that learn object identities and temporal information are proposed to improve multi-player tracking under these conditions.
  • Sports Analysis: Large spatio-temporal datasets, temporal information, subtype labels, and player trajectories are proposed to improve defensive-strategy, action, and team-tactics analysis.
  • Sport-Specific Challenges: Open application challenges include accurate pose recognition under occlusion, cricket run-out and ball-trajectory analysis, badminton camera and illumination robustness, and jersey-number recognition.
  • Ball Analysis: Ball detection and tracking remain open problems because balls are small, fast, and irregularly moving, while current methods are not robust across size, shape, velocity, and environmental conditions.
  • Player Tracking: Player tracking remains difficult because of fast movement, similar jerseys, partial or full occlusions, identity switches, camera viewpoints, and calibration.

9 Conclusion

This review surveys computer-vision methods and applications for sports video analysis, while identifying persistent challenges, research directions, and deployment considerations. It covers detection, tracking, prediction, strategy analysis, event classification, datasets, AI applications, GPU workstations, and embedded platforms.

  • The review covers player and ball detection, tracking, trajectory prediction, team-strategy analysis, and event classification across sports.
  • The review identifies publicly available sport-specific datasets and discusses AI applications, GPU-based workstations, embedded platforms, classical and AI techniques, and future research directions.
  • Algorithms trained for one sport may not transfer to another, motivating datasets containing samples from every sport for fine-tuning.
  • Methods have sport-specific trade-offs: Kalman and particle filters handle ball size, shape, and velocity, while trajectory methods address occlusion but can fail with varying ball size and shape.
  • Data association suits small balls on small courts such as tennis but is less suitable for basketball, soccer, and volleyball challenges.
  • Sports tracking remains difficult because of occlusion, similar appearances, rapid movement, extreme aspect ratios, identity switches, and unreadable jersey numbers.These challenges affect reliable identification of players, referees, and goalkeepers.
  • Quantitative comparison across sports is difficult because experiments, infrastructure, capture devices, and ground-truth databases differ substantially.

Abbreviations:

The paper uses several abbreviations for convolutional neural-network architectures and tracking or sequence-learning methods in sports vision.

  • BEI-CNN denotes Basketball Energy Image-Convolutional Neural Network.
  • Deep-SORT denotes Simple Online Real Time Tracking with Deep Association.
  • Faster-RCNN denotes Faster-Regional with Convolutional Neural Network.
  • GRU-CNN denotes Gated Recurrent Unit-Convolutional Neural Network.
  • Mask R-CNN and R-CNN denote Mask Region-based Convolutional Neural Network and Region-based Convolutional Neural Network, respectively.
Loading 2203.02281v2…