R0: Foundations of the sportsfeatures Analytical Framework

Mohammad Abbas

2026-09-28

Overview

The sportsfeatures analytical framework introduces a rule‑based simulation environment designed to meet the challenges of modern statistical training. The primary purpose of sportsfeatures is to provide a realistic and interpretable environment for learning modern statistical modelling and data science techniques. The framework enables users to explore hierarchical data structures, missing-data mechanisms, feature engineering, predictive modelling, and latent structure discovery within a coherent sporting context.

Why sportsfeatures?

Modern statistical education often relies on small, highly curated datasets that illustrate individual methods in isolation. While useful for teaching specific techniques, such datasets rarely capture the complexities encountered in real analytical settings. Hierarchical structures, missing data, behavioural variability, feature engineering, and latent patterns are frequently absent or heavily simplified. The sportsfeatures framework was designed to address this gap by providing a realistic yet controlled simulation environment for statistical learning. Rather than focusing on a single analytical objective, the framework integrates multiple challenges commonly encountered in applied data science, including multilevel data structures, non constant variance, engineered features, and missing-data mechanisms. Sport and human performance provide a particularly useful context for this purpose. Athletic data naturally contain repeated observations, individual baselines, environmental influences, and physiological variability. These characteristics create a rich setting for exploring modern modelling techniques while maintaining an intuitive connection to real-world processes. The goal of sportsfeatures is not merely to generate synthetic observations but to create an extensible analytical framework. The resulting environment allows users to progress from foundational concepts through prediction, explanation, discovery, and dimensional reduction, forming the basis of the R-series vignette pathway presented throughout this package.

Core Premise

Stable Subject Baselines: Each athlete maintains stable physiological characteristics.

Dynamic Session Evolution: Training sessions vary across environmental, physiological, and personal conditions.

Controlled Realism: Incorporates domain rules, controlled randomness, and a deliberate Missing Not At Random (MNAR) mechanism tied to device usage.

Framework Structure

Phase 1 – Simulation Pipeline

Conceptual idea and physiological hypothesis

Rule‑based simulation engine

Core dataset: 500 observations × 30 variables

Phase 2 – Feature Engineering & Data Enhancement

Enhanced dataset: 500 × 42 variables

Key analytical measures: efficiency, workload normalisation, baselines, variability

Framework Core – Three Analytical Pillars

Pillar 1: Variance Structure – within‑ vs between‑athlete fluctuations Pillar 2: MNAR Missingness – non‑random dropout due to fatigue/injury Pillar 3: Derived Variables – feature engineering transforming metrics into insights

Vignette Research Pathway

R0 – Foundation → R1 – Predict → R2 – Explain → R3 – Discover → R4 – Reduce

Conceptual Overview of the R0 AnalyticalFramework Figure 1.1 Conceptual Overview of the R0 Analytical Framework

As shown in Figure 1.1, the multi-phase structure of the sportsfeatures environment illustrates how rule‑based simulation, data enhancement, and analytical pillars integrate to form the foundation for the R‑series vignettes.

Pillar 1: Variance Structure & Hierarchical Non‑Constant Spread

Human physiology and athletic behaviour are inherently dynamic and non‑linear. When performance metrics are analysed using traditional linear models, non‑constant variance (heteroscedasticity) naturally emerges — a reflection of biological complexity rather than a modelling flaw.

Baseline Illustration

A simple Ordinary Least Squares (OLS) model relating energy expenditure to session volume serves as a baseline to demonstrate non-constant variance.

Calories Burned = β0 + β1(Distance KM) + ε

df_lm <- df_sports %>%
  filter(!is.na(calories_burned), !is.na(distance_km))

model_simple <- lm(calories_burned ~ distance_km, data = df_lm)

Bivariate Scatterplot of Calories Burned vs DistanceKM Figure 1.2 Bivariate Scatterplot of Calories(kcal) Burned vs Distance(km)

Figure 1.2 demonstrates a relatively strong positive trend between total energy expenditure (kcal) and distance covered (km). However, looking closely at the spread of data points, it can be seen that the spread is fanning out, suggesting the presence of non-constant variance (heteroscedasticity).

Figure 1.3 below confirms the presence of non-constant variance. The plot also reveals that as training volume increases, the spread of observations widens — signalling the influence of physiological, behavioural, and environmental factors.

Residual Diagnostics for Simple OLS Model Figure 1.3 Residual Diagnostics for Simple OLS Model

Table 1.1: Simple Linear Regression Model Coefficients

Characteristic Beta 95% CI p-value
(Intercept) -29 -95, 38 0.4
Distance (km) 59 53, 65 <0.001

Adjusted R² 47.6%

While the simple regression model calories_burned ~ distance_km captures a substantial proportion of the variation in calorie expenditure (Adjusted R² = 47.6%), considerable unexplained variation remains. Importantly, the model assumes that all observations are independent and that any remaining variability can be treated as random error. In practice, however, repeated training sessions are recorded from the same athletes. Individual differences in fitness, physiology, training habits, and recovery status introduce structure into the residual variation. Consequently, part of the unexplained variance reflects meaningful athlete-level differences rather than pure noise. To distinguish within-athlete fluctuations from between-athlete differences, the sportsfeatures framework adopts a hierarchical perspective. This transition motivates the mixed-effects modelling approaches explored in later vignettes, where sources of variation can be explicitly partitioned and interpreted.

Hierarchical Variance Decomposition

Because observations are nested within athletes, mixed-effects models provide a natural mechanism for separating within-athlete variability from between-athlete differences.

In the sportsfeatures framework, this non‑constant spread arises from nested sources of variability:

Within‑Athlete vs Between‑Athlete Variation: Repeated sessions within individuals differ from those across athletes with distinct fitness baselines and metabolic efficiencies.

Contextual & Environmental Drivers: Terrain, weather, fatigue, hydration, and readiness states introduce additional layers of variance.

Rather than treating these fluctuations as random noise, sportsfeatures partitions them through hierarchical and mixed‑effects modelling (see R1–R2), enabling structured interpretation of nested dependencies.

This aligns with Figure 1.1, where Variance Structure forms the apex of the framework’s three analytical pillars.

Modelling Implications

Non-constant variance in sportsfeatures is not a defect but a hallmark of realism. Traditional approaches often suppress heteroscedasticity through transformations; here, it is embraced as evidence of biological and behavioural variability across athletes and training contexts.

Recognising this variability has important modelling implications. Statistical methods that account for hierarchical structure, repeated observations, and athlete-specific effects are often more appropriate than approaches that assume complete independence among observations. Consequently, the framework provides a natural setting for exploring mixed-effects models, variance decomposition, and athlete-level heterogeneity in subsequent vignettes.

Within the sportsfeatures framework, variability is treated not merely as noise to be removed but as information to be understood. This perspective underpins the analytical progression from prediction (R1) to explanation (R2), discovery (R3), and dimensional reduction (R4).

Pillar 2: Missing Not At Random (MNAR) & MICE Imputation

Missing data is an inherent feature of real‑world sports telemetry. The sportsfeatures framework deliberately embeds a deterministic Missing Not At Random (MNAR) mechanism to mirror realistic sensor dropout and athlete fatigue patterns.

MICE was selected because it preserves multivariate relationships across mixed physiological metrics while quantifying uncertainty through multiple imputations. Unlike single-value imputation approaches, MICE generates multiple plausible datasets, allowing uncertainty associated with missing values to be propagated into subsequent analyses. This makes it particularly well suited to the sportsfeatures framework, where physiological variables are often correlated and influenced by common underlying processes.

Wearable Device Dropout Mechanics

Whenever device_type == "none", key physiological streams — heart_rate_avg,hydration_status and exhaustion_level — are absent by design. This configuration replicates common wearable data gaps caused by device failure, non‑compliance, or extreme fatigue.

Across the full dataset, this mechanism introduces 97 missing values (≈ 19.4%) per affected variable, creating a controlled environment for testing imputation strategies.

Imputation Workflow

To preserve the multivariate coherence of athletic performance data, missing entries are handled using Multivariate Imputation by Chained Equations (MICE) with predictive mean matching (PMM):

mice_vars <- df_sports %>%
  select(calories_burned, distance_km, duration_min,
         heart_rate_avg, exhaustion_level, hydration_status, device_type)

mice_model <- mice(mice_vars, m = 5, maxit = 20,
                   method = "pmm", seed = 123, quiet = TRUE)

This iterative process models each variable conditionally, drawing plausible replacements that maintain physiological realism and statistical consistency.

Diagnostic Verification: Iterative Chain Convergence

The convergence plot (Figure 1.4) illustrates how the imputation chains stabilize across 20 iterations for key metrics — calories_burned, heart_rate_avg, and hydration_status. Left panels show mean trajectories; right panels show standard deviations. Consistent crossover among independent chains indicates healthy mixing and convergence toward a stationary imputation distribution.

Convergence Diagnostics for Mice Imputation Chains Figure 1.4 MICE Convergence Diagnostics

Interpretation

Pillar 3: Derived Variables & Physiological Constructs

Where Pillar 1 explains variability and Pillar 2 addresses incompleteness, Pillar 3 transforms raw measurements into interpretable physiological constructs.

Derived variables form the dynamic layer of the sportsfeatures framework. While core variables describe observable training outcomes such as distance covered, session duration, or energy expenditure, they provide only a partial view of athlete performance. Derived variables extend these measurements by capturing underlying physiological processes, workload characteristics, and athlete-specific deviations from baseline behaviour.

Within the sportsfeatures framework, feature engineering serves as a bridge between raw observations and analytical insight. By transforming primary measurements into meaningful analytical features, derived variables support prediction, explanation, discovery, and dimensional reduction across the R-series vignettes.

Consequently, derived variables do not simply increase the number of predictors available for analysis; they provide mechanisms through which variability can be understood, modelled, and interpreted.

Table 1.2 summarises the principal derived variables within the framework, along with their hierarchical classification and intended role in modelling.

Some variables capture session-level physiological responses, whereas others represent athlete-level baselines or deviations from those baselines. Together, they provide the analytical bridge between raw observations and higher-level statistical interpretation.

Derived Variable Level Classification Role in Modelling Process
vo2_utilisation Level‑1 (Session) Normalises aerobic demand relative to VO₂_max; quantifies session intensity as a fraction of physiological capacity.
efficiency_index Level‑1 (Session) Captures cardiovascular economy by relating oxygen uptake to heart rate; used to assess aerobic efficiency and readiness.
hr_reserve Level‑1 (Session) Represents net cardiovascular elevation above resting state; proxy for internal load and autonomic strain.
aerobic_efficiency Level‑1 (Session) Indicates aerobic contribution per unit heart rate; evaluates aerobic system utilisation and metabolic economy.
fatigue_per_km Level‑1 (Session) Normalises fatigue accumulation by distance; models workload‑adjusted fatigue dynamics.
athlete_mean_fatigue Level‑2 (Between‑Athlete) Chronic fatigue baseline; separates stable athlete‑level fatigue differences from session‑level variation.
fatigue_deviation Level‑1 (Centered) Within‑athlete centred fatigue; isolates session‑specific deviations from an athlete’s personal fatigue norm.
athlete_mean_distance Level‑2 (Between‑Athlete) Chronic volume baseline; captures habitual training load differences between athletes.
distance_centered Level‑1 (Centered) Within‑athlete centred distance; models session‑specific deviations from typical training volume.
pace_min_km Level‑1 (Session) Represents external speed demand; used to model performance intensity and pacing strategy.
calories_per_km Level‑1 (Session) Normalises metabolic cost by distance; quantifies energetic efficiency and workload economy.
calories_per_min Level‑1 (Session) Represents metabolic burn rate; models internal metabolic intensity independent of distance.

Notes on Multilevel Classification

Note

Evolution from Simulation to Feature Engineering Figure 1.5 Evolution from Simulation to Feature Engineering

As illustrated in Figure 1.5, the sportsfeatures framework evolves through three stages:

The derived variables generated during Stage 3 serve distinct analytical purposes within the framework. Some variables quantify physiological efficiency, others capture workload intensity, chronic athlete baselines, or session-specific deviations from those baselines. Collectively, they provide the mechanisms through which variability can be explored, explained, and modelled.

Derived variables map Figure 1.6 Derived Variables Functional Map

These constructs — such as Heart Rate Reserve, Fatigue Deviation, VO₂ Utilisation, and Efficiency Index — act as mechanism generators. They reveal the underlying processes driving variability in human performance, allowing the framework to decompose, model, and interpret fluctuations rather than treat them as noise.

The derived variables support sportsfeatures’ transition from a dataset to a fully functional framework, thereby enabling hierarchical modelling, interaction profiling, and advanced statistical exploration across vignettes R1 to R4.

What Next? From Framework to Analytical Discovery

The sportsfeatures framework establishes a structured foundation that transforms raw athletic telemetry into biological signals. With individual baseline variance partitioned, missingness appropriately mapped, and 12 feature domains fully established, sportsfeatures evolves from a dataset into an extensible analytical framework.

Framework to analytical discovery Figure 1.7 Framework to Analytical Discovery

This sets the stage for a four-part empirical progression across R1 through R4, as illustrated in Figure 1.7:

Concluding Remarks

R0 establishes the conceptual foundations of the sportsfeatures framework by introducing its simulation architecture, variance structure, missing-data mechanisms, and derived-variable ecosystem. Together, these components create a coherent analytical environment that supports increasingly sophisticated statistical investigations. The subsequent vignettes build upon this foundation, progressing from prediction and explanation to discovery and dimensional reduction within a unified modelling framework.