$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The study included 40 male football players aged 18-24 years (mean = 20.1 ± 1.2), recruited from four amateur and semi-professional clubs within one regional league. All participants had at least three years of competitive experience and were currently training in a structured environment. Data collection was performed in two controlled settings: first, in standardized training-ground drills that simulated tactical pass/drive decision scenarios; second, in prolonged sessions of match simulation from which multimodal sequences were extracted. All procedures were approved by the Institutional Ethics Committee of Beibu Gulf University (Approval No: BG-PE-2023-117), and written informed consent was obtained from all participants prior to data collection. International and institutional ethical standards were adhered to in conducting this investigation. The Beibu Gulf University Institutional Ethics Committee examined and authorized all procedures involving human subjects (Approval No: BG-PE-2023-117). Prior to data collection, each participant provided written informed consent, and all collected data were anonymized before analysis.
This work seeks to automate and enhance the investigation of team sport tactical decision-making through a Transformer-based Multimodal Fusion architecture. It uses a variety of data sources, including video, audio, spatial location, and player metadata to create a robust real-time predictive model. Figure 1 shows the architecture of the proposed method. The proposed work consists of five stages. The stages are described below:

Figure 1. Architecture of the proposed method. Schematic of the model combining multimodal inputs and classification components for tactical decision-making. Please click here to view a larger version of this figure.
Data acquisition and preprocessing
Data from various modalities are gathered from recordings of matches, such as video feeds, audio commentary, GPS/positioning information, and player performance indicators. Each modality undergoes preprocessing in the form of techniques such as frame extraction (for video), noise filtering (for audio), and normalization (for sensor readings), ensuring synchronization at time steps.
Video frames were extracted at 25 fps, downsampled to 224 × 224 resolution. Audio signals were processed with a band-pass filter, ranging from 300 to 3,400 Hz, to eliminate the majority of environmental noise. Positional tracking data from GPS sensors were recorded at 10 Hz and interpolated using linear timestamp-based interpolation, ensuring they align with the 25 Hz video timeline. All inputs were temporally aligned using shared timestamps, and their scales were normalized with z-scoring.
Feature extraction
To gather the spatiotemporal information of video data, 3D CNNs are used. The 3D CNNs handle both spatial and temporal information across video frames, making the model capable of detecting movement patterns, understanding group behavior, and extracting motion-based features. This operation generates a dense feature representation per action sequence, which is used as the visual input to the model. Each video input clip contained 16 consecutive frames sampled at a stride of 2. Random cropping, horizontal flipping, and brightness normalization were applied during training. The 3D CNN was optimized using Adam (lr = 0.0001), batch size 16, and cross-entropy loss.
Embedding and alignment of multimodal inputs
After that, embeddings from different modalities are integrated with the 3D CNN's outputs. A multimodal transformer model architecture is employed, with various encoder branches handling each modality. Cross-attention layers are employed for modality fusion and alignment, enabling the model to capture interdependencies (e.g., between player locations and ball movement, and context from commentary). This fusion allows for an enriched contextual understanding of current tactical choices. Also implemented all embeddings in PyTorch. Video features were projected linearly from 1,024 to 512 dimensions, and audio and positional features were projected to 256 and 128 dimensions, respectively, unified to 512 dimensions before temporal alignment.
Transformer-based multimodal fusion
A Multimodal Transformer is employed to combine the features extracted from all sources. Self-attention mechanisms in the transformer enable the model to Learn interdependencies between modalities (e.g., visual + audio), capture temporal flow and context, and highlight tactically important frames and events. The outcome of this step is a combined representation that captures the tactical intent and situational context of each decision made. The multimodal transformer consisted of six encoder layers, each with eight attention heads. Attention matrices were computed using scaled dot-product attention. Attention and feed-forward layers also included dropout with p = 0.1 to avoid overfitting.
Tactical decision classification and evaluation
Lastly, the fused representation is fed as input to a classification layer or decision analysis block, which is used to forecast the tactical decision type, assess decision quality, and optionally generate explainable insights for coaches through attention heatmaps or activation maps. The output of this stage provides a quantitative and qualitative evaluation of a team or player's decision-making abilities during match play.
The classification head utilised a three-layer MLP with ReLU activation followed by a softmax layer. The model was trained on a 70/15/15 split with 10-fold cross-validation. Metrics used were accuracy, precision, recall, F1-score, latency, and AUC.
A simplified flowchart, summarizing the five stages, provides an immediate visual overview of the sequential process, as shown in Figure 2. This figure concisely represents the complete research pipeline, comprising data acquisition, preprocessing, feature extraction, multimodal fusion, and tactical decision classification, followed by detailed descriptions of each stage.

Figure 2. Overall workflow of the research process. Overview of the research pipeline, from data collection and preprocessing to feature extraction, model training, and evaluation. Please click here to view a larger version of this figure.
Figure 3 displays the overall workflow of the suggested research process. The process is initiated by Stage 1: Data Acquisition and Pre-processing, where multimodal data sources, such as video, audio, positional tracking, and contextual metadata, are gathered and synchronized. During Stage 2: Feature Extraction, 3D Convolutional Neural Networks (3D CNNs) are utilized to extract spatiotemporal information from video streams, resulting in dense visual representations of player actions and tactical flow. Stage 3: Embedding and Alignment of Multimodal Inputs combines audio, video, and positional features into a shared latent space, with temporal alignment and cross-modal alignment. These representations are subsequently aggregated in Stage 4: Transformer-based Multimodal Fusion, where cross-modal and self-attention mechanisms allow the model to learn interdependencies across modalities and derive deeper contextual insights into tactical decisions. Lastly, in Stage 5: Tactical Decision Classification and Evaluation, the combined representations are fed into a classification layer to detect tactical decisions, such as pass, drive, or hold. Meanwhile, performance metrics are employed to estimate decision-making accuracy and response time. This flowchart gives a clear overview of the step-by-step approach, complemented by the elaborate textual descriptions outlined in the following sections.

Figure 3. Sequential processing pipeline. Depicts the step-by-step flow from data acquisition to tactical decision classification and model output. Please click here to view a larger version of this figure.
Data acquisition and preprocessing
To enable high-level tactical decision-making analysis in team sports via a multimodal fusion pipeline using Transformer, a structured and high-fidelity data acquisition and preprocessing setup was formed. The setup involves fusing multiple data modalities, including video recordings, player positional tracking, audio signals from player-coach interactions, and contextual metadata such as match phase, team structure, and possession areas.
The study sample consisted of 40 males aged 20 and above, all of whom were actively involved in regional amateur and semi-professional football leagues. Each of them was recruited from four competitive clubs that had a structured training program. The athletes' mean age was 20.1 ± 1.2 years, with a mean height of 177.5 ± 3.1 cm and a mean weight of 72.6 ± 2.9 kg. They had been playing for their respective clubs for a mean of 6.8 ± 1.5 years. All the participants also had a minimum of three years of competitive experience and were found to be tactically competent.
To ensure statistical power, the sample size was approximated using a 95% level of reliability (Z = 1.96) and a 5% significance level with an expected Intraclass Correlation Coefficient (ICC) of 0.96 and a 95% confidence interval to reflect near-perfect agreement. This is consistent with the objective of validating the reliability and consistency of multimodal decision-making data collected27.
To ensure reproducibility, acquisition settings and technical specifications were standardized across sessions. Video data were recorded at 1,080p resolution and 25 frames per second using GoPro Hero4 cameras with a wide field-of-view setting to capture the full tactical space. Each recording was stabilized and mounted at a fixed height of 2.5 m. Audio signals were captured using an external field microphone at a 44.1 kHz sampling rate (mono) and subsequently filtered using a 300-3,400 Hz band-pass filter to remove environmental noise. Positional tracking data were collected using wearable GPS sensors operating at 10 Hz, with an average manufacturer-reported positional accuracy of ±0.5 m. All modalities were synchronized using shared timestamps and resampled to a common 25 Hz timeline through linear interpolation.
Preprocessing then included the following: video downsampling to 224 × 224 resolution, frame-sampling of 16 consecutive frames with a stride of 2, random cropping, horizontal flipping, brightness normalization, audio normalization, and filtering out noise, and positional and contextual variable z-score normalization.
The preprocessing script is arranged in the following order for reproducibility: (i) frame extraction; (ii) audio filtering; (iii) GPS interpolation; (iv) modality resampling to 25 Hz; (v) z-score normalization; and (vi) integrity tests for corruption, occlusion, and dropout. Batch-processing scripts are used to automatically complete each stage.
To keep data reliable, the criteria for quality control and exclusion included: i) sequences with >20% player occlusion, ii) GPS dropout exceeding 200 ms, iii) corrupted or clipped audio, and iv) severe motion blur or desynchronization between modalities. A total of 17 sequences (3.4%) were excluded for these conditions. Table 1 shows the dataset description.
| Attributes | Details |
| Sport | Football (Soccer) |
| Participants | 40 male football players |
| Level of Play | Federated football players with ≥5 years of experience |
| Test Type | Passing and driving (ball control) decision-making and execution test |
| Test Setting | Standardized football pitch setup, marked zones, fixed positions |
| Measurement Tools | GoPro Hero4 cameras, Kinovea software, stopwatch timing |
| Data Collected | Execution Time (ET), Decision-Making (DM) Accuracy, Number of Correct Actions |
| Test Procedure | Players responded to visual stimuli from coaches and executed passes or drives |
| Trial Repetitions | 10 trials per player (5 passes, 5 drives) |
| Evaluation | Experts rated DM; ET measured using timestamps/video |
| Outcome Variables | Reaction time, correct decision count, technical execution score |
Table 1: Dataset description. Details participant demographics, testing procedures, measurement tools, and outcome variables.
In the present study, two related but distinct datasets were used. The first dataset comprised 40 players who participated in controlled pass/drive decision-making tasks, primarily used to establish baseline reliability and validate the basic decision measurement procedure. The second dataset includes 500 multimodal tactical sequences collected from extended match simulations and real-game recordings. These sequences constituted the main training and testing dataset for the transformer-based fusion model, allowing evaluation under more ecologically valid conditions.
Although the dataset comprises 500 multimodal sequences, stratified sampling ensured balanced representation across tactical actions (Pass, Drive, Hold). Data augmentation approaches, such as temporal cropping and sequence shuffling, were also utilized to enhance variability and mitigate model bias during training.
It is worth noting that the 40-player controlled dataset was used in the initial model validation and testing of reliability, while the larger collection of 500 multimodal sequences was compiled from prolonged match recordings and organized training sessions to assess the scalability and ecological validity of the framework. Baseline reliability was established using the smaller dataset, while the larger dataset facilitated large-scale model training and testing.
Dataset description
The experiment recruited 40 male football players, each with at least five years of federated playing experience, to ensure that the sample population had basic technical and tactical competencies. The data gathering took place in a standardized football pitch arrangement where predetermined zones and playing positions were assigned to ensure identical testing conditions. During experimentation, participants had to respond to visual cues from a coach, either to pass or drive the ball, replicating tactical decisions made in an actual game. Each participant completed a total of 10 trials, five passing and five driving. These movements were captured with GoPro Hero4 cameras, and performance data were measured using Kinovea video analysis software alongside manual stopwatch time to capture execution time (ET). The main data source gathered was: execution time, the quality of the DM as evaluated by experienced coaches, and the quantity of correct actions executed. These factors provided a quantitative foundation for analyzing how quickly and effectively players were in making and carrying out tactical decisions under controlled stimulus-response conditions. The dataset produced is a structured and labelled resource well-suited to the task of creating benchmarks or training intelligent systems with the goal of modelling tactical thinking and motor responses in team sports contexts.
The sample is restricted to young male players from one regional league, which may limit its generalizability across age groups, female athletes, or other cultures. Future studies will attempt to draw more representative and diverse samples.
Another limitation is the sample size of 40 young male players from a single regional league. While the homogeneous sample contributed to internal validity and reduced variation in performance due to age, gender, and tactical background differences, it also limits the generalizability of findings to a wider population. The tactical patterns the model learns may thus reflect region-specific or gender-specific playing styles rather than universal decision-making behavior. For this study, stratified sampling and data augmentation were employed to enhance the variability of the dataset at hand; however, larger-scale research studies involving female athletes, youth players, professional players, and participants across multiple tactical cultures are needed to confirm the robustness of the model across diverse football environments. These results demonstrate strong promise but need further validation on more diverse and larger datasets before wider generalization can be claimed.
To reduce subjectivity, three professional football coaches independently annotated the tactical choices. Inter-rater reliability was estimated using the intraclass correlation coefficient (ICC = 0.96, CI 0.94-0.98), indicating near-perfect agreement. Consensus discussion resolved disagreement. For multimodal data, preprocessing entailed temporal synchronization between modalities (25 fps video and 10 Hz sensor data) and z-score normalization of features to allow comparability across modalities.
Inter-rater reliability was quantified using a two-way random-effects intraclass correlation coefficient (ICC [2,3]), as it is suited for assessing absolute agreement between multiple raters. The ICC of 0.96 (95% CI: 0.94-0.98) indicates near-perfect agreement according to the commonly used thresholds and further establishes that the tactical decision labels used for training were highly dependable. Reliability was computed across all 500 multimodal sequences, ensuring that both controlled and match-simulation contexts were duly and consistently annotated.
Annotation procedure and labeling protocol
All multimodal sequences were annotated using a structured, three-stage protocol. First, three licensed football coaches independently reviewed each video-positional sequence using a timestamp-synchronized annotation interface (Kinovea + custom tagging spreadsheet). Annotators labeled the tactical decision category (Pass, Drive, Hold), contextual cues (pressure zone, passing lane availability), and event timestamps. Second, a disagreement-detection script automatically identified cases of ambiguity or conflict, which were then reviewed in a joint consensus meeting. Third, to ensure consistency in event boundary marking and temporal alignment across modalities, a final pass was conducted by an independent senior analyst. This process, for the purposes of later model training, ensured that every annotation was traceable and of the best quality.
Preprocessing
Whereas the present study was conducted with 500 multimodal sequences, the preprocessing pipeline was, by design, set up for large-scale datasets. All modalities underwent automated scripts that, in parallel, batch-extracted video frames, resampled audio, interpolated positional data, and applied z-score normalization. The presence of modular architecture allows each modality to be independently preprocessed on separate GPU/CPU threads. This enables the high-volume ingestion of match footage. This ensures that the workflow can easily be scaled to full-season datasets without manual intervention and makes the framework suitable for real-world deployment in professional analytics environments.
To properly examine tactical decision-making in team sports, raw data from different modalities-e.g., video clips, audio signals, and positional recordings-need to be pre-processed and aligned. The multimodal input illustration is
(1)
Here, V is denoted as series of video characteristics extracted from CNN encoder.
(2)
Here, A is an audio characteristic is encoded by wave2vec.
(3)
P represents the coordinate position of a player at time ti.
To deal with asynchronous sampling rates, each modality is resampled or interpolated to a shared timeline:
(4)
Here
, is a synchronized characteristic ft , is rate of combined frame, Xm is an initial indication.
For uniform scale across modalities, all feature vectors are z-score normalized:
(5)
μX and σX denote the mean and the standard deviation of the feature dimension in the training set.
While the primary concerns are pass/drive decisions, audio channels (e.g., player/coach communication, environmental game cues) were added to account for actual-match situations where speech signals affect tactical reaction. Hyperparameters (e.g., learning rate, epochs) were chosen through grid search and existing evaluations in multimodal learning, striking a balance between convergence stability and computational efficiency.
The hyperparameters of the suggested framework-learning rate, batch size, and training epochs-were chosen through a well-planned systematic grid search to achieve both convergence stability and computational efficiency. The choice of a learning rate of 0.0001 with the Adam optimizer was found to provide stable updates of the gradient without oscillations, while batch sizes were optimized to balance GPU memory capacity against training throughput. The setting of 50-100 epochs was used to permit adequate learning without overfitting, as confirmed by validation loss tracking and 10-fold cross-validation. These values were not set arbitrarily but were guided by earlier testing in multimodal learning and adjusted by experimental trials on the current dataset. This rigorous process guaranteed that the selected hyperparameters facilitated stable model training and reproducible improvements in performance over baseline approaches.
Although the classification task focuses on pass/drive/hold decisions, audio information was intentionally included since real tactical behavior in football is strongly influenced by player-coach communication, teammate calling cues, pressing signals, and environmental sounds (e.g., ball contact, opponent approach). These cues frequently precede or co-occur with decision-making and form part of the natural perceptual environment of football players. Preliminary ablation experiments showed that excluding audio reduced recall by 3.8%, which indicated that even subtle acoustic cues contribute to context awareness. For this reason, audio was retained to preserve ecological validity and improve the model's ability to generalize to real match conditions.
Indeed, in further support of these hyperparameters, an internal sensitivity analysis was conducted by systematically varying the learning rate from 1e-5 to 5e-4, the batch size from 8 to 32, the number of transformer layers from 4 to 8, and the number of attention heads from 4 to 12. The chosen configuration, with a learning rate of 1e-4, a batch size of 16, and a 6-layer encoder with 8 attention heads, continuously yielded stable convergence with the lowest validation loss, minimizing overfitting across 10-fold cross-validation. Larger batch sizes resulted in gradient instability, while smaller batches increased the noise and delayed the convergence. Similarly, transformer models with more than 8 heads did not show any performance increase, but they did considerably increase computation time. The obtained results justify the final settings of hyperparameters used in this study.
Feature extraction using 3D convolutional neural network
In the proposed Transformer-Based Multimodal Fusion framework for team sport tactical decision-making, the 3D Convolutional Neural Network (3D CNN) is responsible for extracting spatial-temporal features from raw video clips.
In contrast to 2D CNNs, which consider only spatial dimensions (width W and height H), 3D CNNs also account for the temporal dimension R, enabling them to detect motion dynamics and sequential patterns that are crucial for understanding team behaviors, such as driving and passing in football.
The 3D-CNN28 is first proposed to combine the spatial and temporal information of video data. A 3D kernel is used in a 3D convolution algorithm to a cube made up of frames with part of the columns stacked. The 3D convolution can be written in Eqn (6)
(6)
Where:
is the output feature at location (x, y, z) on the jth feature map of the lth layer.
is the weight of the kernel at location (h,w,r) from the mth input feature map.
is the input at spatial-temporal offset from the preceding layer. blj is the bias. f(.) is the activation function (usually ReLU). M is the number of input feature maps from the preceding layer. This equation allows the 3D CNN to combine the spatial (height and width) and temporal (frame sequence) information in parallel.

Figure 4. 3D CNN processing architecture. Illustrates how spatiotemporal features are extracted from video sequences using 3D convolutional operations. Please click here to view a larger version of this figure.
Figure 4 shows the working process of a 3D CNN. The pipeline starts with the input phase, where unprocessed video streams of a football game are fused with organized data, including player locations, game statistics, and contextual metadata. This multimodal information preserves both visual dynamics and situational context, which are critical for team tactics analysis. Then, the video stream is routed through a 3D Convolutional Neural Network (3D CNN). In this case, a series of 3D convolutional layers is used on the input frames. The layers convolve along the spatial dimensions (height and width) and the temporal dimension (frames), enabling the network to learn motion patterns, player interactions, and sequence-based tactics, such as passes or drives. Subsequently, a feature cube is formed that has dense spatiotemporal features. This cube is subjected to pooling operations to reduce the dimension while preserving significant features. Upon pooling, the feature volume is flattened into a 1D vector, providing a concise representation of the tactical situation. This feature vector is subsequently input to a fully connected Deep Neural Network (DNN). The DNN is the decision-making block, which processes the fused and extracted features to make tactical decisions, such as passing, shooting, or holding. The output of the network is the final tactical decision taken by the system based on the actual game situation in real time as well as learned patterns of strategy.
Embedding and alignment of multimodal inputs
After the feature extraction, the three features, such as audio, video, and statistical position of the player, are given as an input modality.
To input these into a transformer and then feed each of them through a learnable embedding function:
(7)
where f3DCNN learns spatiotemporal features of each time slice Vt and maps it to a d-dimensional latent space.
(8)
where WP and bP are learnable weights.
(9)
All embeddings are aligned by the temporal dimension T. If the modalities have disparate frequencies, it uses interpolation or down sampling:
(10)
Now, all modalities are synchronized as:
(11)
In order to preserve temporal order in the transformer, also introduce positional encodings to every vector:
(12)
Where
represents the default sinusoidal or learnable positional encoding.
Each modality is operationally encoded using a PyTorch linear layer, concatenated into a 512-d vector, and then fed into the transformer input sequence following positional encoding. This guarantees that a unified embedding is prepared for cross-attention at each time step.
The Final tensor is
(13)
is fed into the Transformer Encoder, where the self-attention and cross-modal attention mechanisms enable the model to reason jointly across all modalities for tactical decision-making.
Transformer-based multimodal fusion
In the proposed approach, the low-rank matrix factorization method is employed to integrate the acquired multimodal data vectors. The high-dimensional issue that occurs when tensors are fused directly is addressed by this method, which uses a low-rank decomposition factor for tensor fusion29. Each modality is initially expressed as a vector, and dimensionality reduction that preserves significant interactions is achieved using low-rank decomposition. The fusion tensor ZM among M modalities is built as:
(14)
Here:
is the m-th modality's low-rank factor,
is the outer product, and r is the decomposition rank.
The fusion vector h is then calculated as:
(15)
When two modalities are utilized alone, the reduced fusion is:
(16)
This produces a joint multimodal representation vector h that models cross-modal interactions at a latent level.
Following the acquisition of the low-rank fused vector h, the Transformer module models dependencies and aligns semantic representations between modalities. A cross-modal attention is specially designed to enable one modality to pay attention to another:
(17)
Where Qm is the modality m query, Kn,Vn are the key and value vectors in modality n, dk is the dimension of the key vectors. This setup facilitates modality-specific attention flows, allowing each category of input to be processed relative to others (e.g., synchronizing player motion with game context and video indicators).
The cross-attended vectors are stacked and classified through a Multilayer Perceptron (MLP). The output is a tactical decision label (e.g., pass, drive, hold), cross-entropy loss optimized for end-to-end training.

Figure 5. Transformer-based multimodal fusion architecture. Shows how visual and sensor modalities are integrated via a transformer encoder to enhance feature representation. Please click here to view a larger version of this figure.
The multimodal fusion block, as demonstrated in Figure 5, combines visual and positional inputs. It is a necessary design since tactical decisions require spatial awareness and movement dynamics, which cannot be represented by a single modality.
Figure 5 shows the working structure of Transformer-based Multimodal Fusion. This design combines diverse data video, spatial coordinates, and game contextual data into a single channel to predict tactical actions such as passing, driving, or holding. The approach begins with inputting videos, which are vetted using a 3D Convolutional Neural Network (3D CNN). This system identifies both spatial and temporal structures in video, enabling it to capture dynamic player activity and communication over time. Meanwhile, contextual and positional information are conveyed by embedding layers, transforming structured inputs into dense vector representations. The contextual, positional, and video streams can then have their outputs fused by a Low-Rank Fusion module. The approach effectively obtains cross-modal correlations through fusion space decomposition, minimizing computational complexity while preserving essential relationships among modalities. Finally, the fused multimodal representation is processed by a Transformer Encoder with cross-modal attention. This procedure enables the model to focus on the right features across modalities and time steps, enhancing its understanding of strategic game behavior. Finally, the encoded output is passed through a concatenation and a Multi-Layer Perceptron (MLP) block that maps the learned representation to some tactical decisions. The model output classifies player actions such as "pass," "drive," or "hold" to enable intelligent, data-driven examination of team tactics.
In the suggested Transformer-Based Multimodal Fusion tactical analysis system for team sports, the last phase is tactical decision classification and assessment, which converts learned multimodal representations to decision-actionable labels. Having fused video, positional, and contextual features by low-rank fusion and fine-tuned them using cross-modal attention in the Transformer encoder, the resulting feature vectors are fed into a Multi-Layer Perceptron (MLP) classifier. Every decision class is represented with a softmax activation function, yielding probability scores for all possible actions. The most probable class is chosen as the model's predicted decision. The model is trained on a cross-entropy loss function, which rewards correct predictions and causes the network to learn to differentiate between similar differences in tactical situations. Additionally, time-domain evaluation could be considered in order to confirm responsiveness in real-time or near-real-time game situations, ensuring that the classifier is accurate and also not wasteful of time under game-like time constraints. Table 2 shows the Model Architecture and Hyperparameters.
| Component | Specification |
| Transformer Encoder Layers | 6 |
| Attention Heads | 8 |
| Hidden Dimension | 512 |
| Feed-Forward Layer Size | 2048 |
| Embedding Dimensions | Video: 1024 → 512, Audio: 256 → 512, Positional: 128 → 512 |
| Positional Encoding | Learnable |
| Fusion Method | Low-Rank Multimodal Fusion (rank r = 16) |
| Optimizer | Adam |
| Learning Rate | 0.0001 |
| Batch Size | 16 |
| Training Epochs | 50–100 |
| Dropout | 0.1 |
| Loss Function | Cross-Entropy |
Table 2: Model architecture and hyperparameters. Lists training settings and key components of the proposed transformer-based multimodal fusion model.
For further clarity and reproducibility, the multimodal transformer was implemented with a 6-layer encoder stack, each equipped with 8 attention heads and a hidden dimension of 512, following typical transformer settings for mid-scale multimodal tasks. Each modality passes through an embedding block: video by a 3D-CNN encoder output, 1024-dim; audio by a Wave2Vec-based encoder, 256-dim; and positional plus contextual features by a 2-layer MLP encoder, 128-dim. All these embeddings were projected to a shared 512-d latent space through learnable linear transformations. To establish temporal alignment, all modalities are first interpolated onto a single frequency of 25 Hz, followed by the application of learned positional encodings to preserve temporal ordering. Cross-modal alignment is done using the cross-attention layers, which explicitly learn correlations between modalities before being fed into the self-attention encoder stack. These architectural choices were made to balance model capacity with computational efficiency, thereby ensuring stable convergence on a moderately sized dataset.
In practical terms, the proposed Transformer-Based Multimodal Fusion framework is most suitable for scenarios where synchronized multimodal data are available, such as structured training sessions, controlled match simulations, or professional environments equipped with consistent video feeds and positional tracking systems. The four major inputs needed for the method are (i) high-resolution video, (ii) positional tracking data (GPS/LPS), (iii) audio streams, and (iv) contextual metadata-e.g., possession, pressure zones, match phase. To achieve reliable performance, these modalities must be time-aligned with minimal drift (video ≥ 25 fps and positional sensors ≥ 10 Hz), and should maintain stable viewpoints and minimal signal noise. Those settings that fall short of providing these synchronized inputs, for instance, amateur matches with irregular camera angles, live broadcast feeds featuring latency/compression artifacts, or an environment lacking positional tracking infrastructure, will be less relevant to the framework. Moreover, the current system does best in semi-controlled conditions but may be substantially less robust in completely uncontrolled live-match situations, where lighting conditions, occlusion, and dynamic crowd noise affect the quality of data acquisition. These constraints outline the practical boundaries within which this system can be deployed and emphasize that consistent capture of data in multiple modes is essential in applying the proposed model.
Modifications and troubleshooting
In practical implementation, several issues can arise throughout the multimodal pipeline that may reduce model stability or overall performance. One common problem is temporal misalignment between modalities, where video recorded at 25 fps and positional data captured at 10 Hz fall out of sync, leading to mismatched cues that destabilize attention maps and reduce recall. This can be resolved by applying timestamp-based resampling, linear interpolation, and automatically removing sequences with excessive drift. Audio channels may also suffer from noise caused by wind, crowd sounds, or microphone clipping, which results in the transformer ignoring audio tokens and slightly lowering precision. Applying band-pass filtering, spectral gating, and substituting severely corrupted segments with zero vectors helps maintain audio embedding stability. Visual quality issues, such as player occlusions or motion blur, can weaken the 3D-CNN's spatiotemporal representations and increase pass-drive confusion. This is mitigated by excluding heavily occluded sequences and employing augmentation strategies, including random cropping, brightness normalization, and motion-blur simulation. Transformer instability may occur if embedding scales differ across modalities, producing oscillating validation loss or gradient spikes. Ensuring consistent z-score normalization, using a warm-up learning rate schedule, applying gradient clipping, and increasing dropout can restore stable training. Low-rank fusion can also become problematic when the fusion rank is improperly set, causing distorted cross-modal relationships and class-wise imbalance. Maintaining a moderate rank for small datasets, increasing the rank for larger datasets, and applying L2 regularization help preserve balanced fusion. Computational bottlenecks may arise due to large video tensors that cause GPU memory overflow or slow training. These issues can be addressed through mixed-precision (FP16) training, reducing input resolution during early experimentation, or caching precomputed 3D-CNN features. Finally, the natural class imbalance in tactical decisions, especially for "Drive," may reduce recall and cause fluctuations in AUC; applying class-balanced loss, oversampling minority actions, and enriching contextual metadata can substantially improve balance and decision discrimination. This consolidated troubleshooting guidance ensures that researchers and practitioners can maintain stable training conditions, improve reproducibility, and adapt the framework effectively across various datasets and environments.