df_lm <- df_sports %>%
filter(!is.na(calories_burned), !is.na(distance_km))
model_simple <- lm(calories_burned ~ distance_km, data = df_lm)2026-09-28
The sportsfeatures analytical framework introduces a rule‑based simulation environment designed to meet the challenges of modern statistical training. The primary purpose of sportsfeatures is to provide a realistic and interpretable environment for learning modern statistical modelling and data science techniques. The framework enables users to explore hierarchical data structures, missing-data mechanisms, feature engineering, predictive modelling, and latent structure discovery within a coherent sporting context.
Modern statistical education often relies on small, highly curated datasets that illustrate individual methods in isolation. While useful for teaching specific techniques, such datasets rarely capture the complexities encountered in real analytical settings. Hierarchical structures, missing data, behavioural variability, feature engineering, and latent patterns are frequently absent or heavily simplified. The sportsfeatures framework was designed to address this gap by providing a realistic yet controlled simulation environment for statistical learning. Rather than focusing on a single analytical objective, the framework integrates multiple challenges commonly encountered in applied data science, including multilevel data structures, non constant variance, engineered features, and missing-data mechanisms. Sport and human performance provide a particularly useful context for this purpose. Athletic data naturally contain repeated observations, individual baselines, environmental influences, and physiological variability. These characteristics create a rich setting for exploring modern modelling techniques while maintaining an intuitive connection to real-world processes. The goal of sportsfeatures is not merely to generate synthetic observations but to create an extensible analytical framework. The resulting environment allows users to progress from foundational concepts through prediction, explanation, discovery, and dimensional reduction, forming the basis of the R-series vignette pathway presented throughout this package.
Stable Subject Baselines: Each athlete maintains stable physiological characteristics.
Dynamic Session Evolution: Training sessions vary across environmental, physiological, and personal conditions.
Controlled Realism: Incorporates domain rules, controlled randomness, and a deliberate Missing Not At Random (MNAR) mechanism tied to device usage.
Phase 1 – Simulation Pipeline
Conceptual idea and physiological hypothesis
Rule‑based simulation engine
Core dataset: 500 observations × 30 variables
Phase 2 – Feature Engineering & Data Enhancement
Enhanced dataset: 500 × 42 variables
Key analytical measures: efficiency, workload normalisation, baselines, variability
Framework Core – Three Analytical Pillars
Pillar 1: Variance Structure – within‑ vs between‑athlete fluctuations Pillar 2: MNAR Missingness – non‑random dropout due to fatigue/injury Pillar 3: Derived Variables – feature engineering transforming metrics into insights
R0 – Foundation → R1 – Predict → R2 – Explain → R3 – Discover → R4 – Reduce
Figure 1.1 Conceptual Overview of the R0 Analytical Framework
As shown in Figure 1.1, the multi-phase structure of the sportsfeatures environment illustrates how rule‑based simulation, data enhancement, and analytical pillars integrate to form the foundation for the R‑series vignettes.
Human physiology and athletic behaviour are inherently dynamic and non‑linear. When performance metrics are analysed using traditional linear models, non‑constant variance (heteroscedasticity) naturally emerges — a reflection of biological complexity rather than a modelling flaw.
A simple Ordinary Least Squares (OLS) model relating energy expenditure to session volume serves as a baseline to demonstrate non-constant variance.
Calories Burned = β0 + β1(Distance KM) + ε
df_lm <- df_sports %>%
filter(!is.na(calories_burned), !is.na(distance_km))
model_simple <- lm(calories_burned ~ distance_km, data = df_lm) Figure 1.2 Bivariate Scatterplot of Calories(kcal) Burned vs Distance(km)
Figure 1.2 demonstrates a relatively strong positive trend between total energy expenditure (kcal) and distance covered (km). However, looking closely at the spread of data points, it can be seen that the spread is fanning out, suggesting the presence of non-constant variance (heteroscedasticity).
Figure 1.3 below confirms the presence of non-constant variance. The plot also reveals that as training volume increases, the spread of observations widens — signalling the influence of physiological, behavioural, and environmental factors.
Figure 1.3 Residual Diagnostics for Simple OLS Model
| Characteristic | Beta | 95% CI | p-value |
|---|---|---|---|
| (Intercept) | -29 | -95, 38 | 0.4 |
| Distance (km) | 59 | 53, 65 | <0.001 |
Adjusted R² 47.6%
While the simple regression model calories_burned ~ distance_km captures a substantial proportion of the variation in calorie expenditure (Adjusted R² = 47.6%), considerable unexplained variation remains. Importantly, the model assumes that all observations are independent and that any remaining variability can be treated as random error. In practice, however, repeated training sessions are recorded from the same athletes. Individual differences in fitness, physiology, training habits, and recovery status introduce structure into the residual variation. Consequently, part of the unexplained variance reflects meaningful athlete-level differences rather than pure noise. To distinguish within-athlete fluctuations from between-athlete differences, the sportsfeatures framework adopts a hierarchical perspective. This transition motivates the mixed-effects modelling approaches explored in later vignettes, where sources of variation can be explicitly partitioned and interpreted.
Because observations are nested within athletes, mixed-effects models provide a natural mechanism for separating within-athlete variability from between-athlete differences.
In the sportsfeatures framework, this non‑constant spread arises from nested sources of variability:
Within‑Athlete vs Between‑Athlete Variation: Repeated sessions within individuals differ from those across athletes with distinct fitness baselines and metabolic efficiencies.
Contextual & Environmental Drivers: Terrain, weather, fatigue, hydration, and readiness states introduce additional layers of variance.
Rather than treating these fluctuations as random noise, sportsfeatures partitions them through hierarchical and mixed‑effects modelling (see R1–R2), enabling structured interpretation of nested dependencies.
This aligns with Figure 1.1, where Variance Structure forms the apex of the framework’s three analytical pillars.
Non-constant variance in sportsfeatures is not a defect but a hallmark of realism. Traditional approaches often suppress heteroscedasticity through transformations; here, it is embraced as evidence of biological and behavioural variability across athletes and training contexts.
Recognising this variability has important modelling implications. Statistical methods that account for hierarchical structure, repeated observations, and athlete-specific effects are often more appropriate than approaches that assume complete independence among observations. Consequently, the framework provides a natural setting for exploring mixed-effects models, variance decomposition, and athlete-level heterogeneity in subsequent vignettes.
Within the sportsfeatures framework, variability is treated not merely as noise to be removed but as information to be understood. This perspective underpins the analytical progression from prediction (R1) to explanation (R2), discovery (R3), and dimensional reduction (R4).
Missing data is an inherent feature of real‑world sports telemetry. The sportsfeatures framework deliberately embeds a deterministic Missing Not At Random (MNAR) mechanism to mirror realistic sensor dropout and athlete fatigue patterns.
MICE was selected because it preserves multivariate relationships across mixed physiological metrics while quantifying uncertainty through multiple imputations. Unlike single-value imputation approaches, MICE generates multiple plausible datasets, allowing uncertainty associated with missing values to be propagated into subsequent analyses. This makes it particularly well suited to the sportsfeatures framework, where physiological variables are often correlated and influenced by common underlying processes.
Whenever device_type == "none", key physiological streams — heart_rate_avg,hydration_status and exhaustion_level — are absent by design. This configuration replicates common wearable data gaps caused by device failure, non‑compliance, or extreme fatigue.
Across the full dataset, this mechanism introduces 97 missing values (≈ 19.4%) per affected variable, creating a controlled environment for testing imputation strategies.
To preserve the multivariate coherence of athletic performance data, missing entries are handled using Multivariate Imputation by Chained Equations (MICE) with predictive mean matching (PMM):
mice_vars <- df_sports %>%
select(calories_burned, distance_km, duration_min,
heart_rate_avg, exhaustion_level, hydration_status, device_type)
mice_model <- mice(mice_vars, m = 5, maxit = 20,
method = "pmm", seed = 123, quiet = TRUE)This iterative process models each variable conditionally, drawing plausible replacements that maintain physiological realism and statistical consistency.
The convergence plot (Figure 1.4) illustrates how the imputation chains stabilize across 20 iterations for key metrics — calories_burned, heart_rate_avg, and hydration_status. Left panels show mean trajectories; right panels show standard deviations. Consistent crossover among independent chains indicates healthy mixing and convergence toward a stationary imputation distribution.
Figure 1.4 MICE Convergence Diagnostics
Mean & Variance Stabilization: The imputed streams interweave smoothly, confirming robust chain mixing and an absence of systemic drift.
Physiological Coherence: The completed data vectors remain statistically valid and biologically plausible, ensuring unbiased downstream modelling.
Framework Integration: This pillar complements Variance Structure (Pillar 1) by addressing data incompleteness — reinforcing the realism of the sportsfeatures simulation pipeline.
Where Pillar 1 explains variability and Pillar 2 addresses incompleteness, Pillar 3 transforms raw measurements into interpretable physiological constructs.
Derived variables form the dynamic layer of the sportsfeatures framework. While core variables describe observable training outcomes such as distance covered, session duration, or energy expenditure, they provide only a partial view of athlete performance. Derived variables extend these measurements by capturing underlying physiological processes, workload characteristics, and athlete-specific deviations from baseline behaviour.
Within the sportsfeatures framework, feature engineering serves as a bridge between raw observations and analytical insight. By transforming primary measurements into meaningful analytical features, derived variables support prediction, explanation, discovery, and dimensional reduction across the R-series vignettes.
Consequently, derived variables do not simply increase the number of predictors available for analysis; they provide mechanisms through which variability can be understood, modelled, and interpreted.
Some variables capture session-level physiological responses, whereas others represent athlete-level baselines or deviations from those baselines. Together, they provide the analytical bridge between raw observations and higher-level statistical interpretation.
| Derived Variable | Level Classification | Role in Modelling Process |
|---|---|---|
| vo2_utilisation | Level‑1 (Session) | Normalises aerobic demand relative to VO₂_max; quantifies session intensity as a fraction of physiological capacity. |
| efficiency_index | Level‑1 (Session) | Captures cardiovascular economy by relating oxygen uptake to heart rate; used to assess aerobic efficiency and readiness. |
| hr_reserve | Level‑1 (Session) | Represents net cardiovascular elevation above resting state; proxy for internal load and autonomic strain. |
| aerobic_efficiency | Level‑1 (Session) | Indicates aerobic contribution per unit heart rate; evaluates aerobic system utilisation and metabolic economy. |
| fatigue_per_km | Level‑1 (Session) | Normalises fatigue accumulation by distance; models workload‑adjusted fatigue dynamics. |
| athlete_mean_fatigue | Level‑2 (Between‑Athlete) | Chronic fatigue baseline; separates stable athlete‑level fatigue differences from session‑level variation. |
| fatigue_deviation | Level‑1 (Centered) | Within‑athlete centred fatigue; isolates session‑specific deviations from an athlete’s personal fatigue norm. |
| athlete_mean_distance | Level‑2 (Between‑Athlete) | Chronic volume baseline; captures habitual training load differences between athletes. |
| distance_centered | Level‑1 (Centered) | Within‑athlete centred distance; models session‑specific deviations from typical training volume. |
| pace_min_km | Level‑1 (Session) | Represents external speed demand; used to model performance intensity and pacing strategy. |
| calories_per_km | Level‑1 (Session) | Normalises metabolic cost by distance; quantifies energetic efficiency and workload economy. |
| calories_per_min | Level‑1 (Session) | Represents metabolic burn rate; models internal metabolic intensity independent of distance. |
Note
- Level‑1 (Session level): varies within athlete across sessions
- Level‑2 (Athlete level): stable between athletes; group‑level baselines
- Level‑1 (Centered): session‑level deviations from Level‑2 baselines
Figure 1.5 Evolution from Simulation to Feature Engineering
As illustrated in Figure 1.5, the sportsfeatures framework evolves through three stages:
Simulation Dataset (Stage 1): establishes baseline synthetic behaviour.
Core Expansion (Stage 2): introduces physiological depth through additional subject-level characteristics.
Feature Engineering (Stage 3): generates twelve derived variables that transform raw measurements into analytical constructs.
The derived variables generated during Stage 3 serve distinct analytical purposes within the framework. Some variables quantify physiological efficiency, others capture workload intensity, chronic athlete baselines, or session-specific deviations from those baselines. Collectively, they provide the mechanisms through which variability can be explored, explained, and modelled.
Figure 1.6 Derived Variables Functional Map
These constructs — such as Heart Rate Reserve, Fatigue Deviation, VO₂ Utilisation, and Efficiency Index — act as mechanism generators. They reveal the underlying processes driving variability in human performance, allowing the framework to decompose, model, and interpret fluctuations rather than treat them as noise.
The derived variables support sportsfeatures’ transition from a dataset to a fully functional framework, thereby enabling hierarchical modelling, interaction profiling, and advanced statistical exploration across vignettes R1 to R4.
The sportsfeatures framework establishes a structured foundation that transforms raw athletic telemetry into biological signals. With individual baseline variance partitioned, missingness appropriately mapped, and 12 feature domains fully established, sportsfeatures evolves from a dataset into an extensible analytical framework.
Figure 1.7 Framework to Analytical Discovery
This sets the stage for a four-part empirical progression across R1 through R4, as illustrated in Figure 1.7:
R1 (Predict): Evaluating mechanical efficiency against athlete readiness using predictive multilevel frameworks.
R2 (Explain): Deconstructing cardiorespiratory cost and partitioning within- versus between-athlete variance.
R3 (Discover): Uncovering hidden physiological taxonomies through unsupervised profiling.
R4 (Reduce): Isolating key latent dimensions across dynamic training variables using factor reduction.
R0 establishes the conceptual foundations of the sportsfeatures framework by introducing its simulation architecture, variance structure, missing-data mechanisms, and derived-variable ecosystem. Together, these components create a coherent analytical environment that supports increasingly sophisticated statistical investigations. The subsequent vignettes build upon this foundation, progressing from prediction and explanation to discovery and dimensional reduction within a unified modelling framework.