TY - GEN
T1 - The Data-Schema Bottleneck
T2 - 9th International Workshop on Artificial Intelligence Techniques for Data Management, aiDM 2026
AU - Debnath, Tanmoy
AU - Debnath, Sourabhi
AU - Narbutt, Miroslaw
AU - Bhattacharya, Maumita
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/7/16
Y1 - 2026/7/16
N2 - Effective management of Low Earth Orbit (LEO) satellite networks depends on data pipelines capable of predicting link-layer quality metrics (e.g., round-trip time (RTT), throughput (TP), and jitter etc.) at timescales suitable for real-time distributed data routing, handover, and congestion management. This study systematically benchmarks sixteen deep learning (DL) architectures representing four inductive-bias families to evaluate two formally stated hypotheses using 30 days of real Starlink telemetry comprising 417 million observations. Firstly, the Spectral Alignment Hypothesis (Research Question (RQ) 1) investigates whether architectures possessing inductive biases that explicitly decompose the quasi-periodic orbital structure of LEO dynamics systematically outperform those processing the telemetry as generic sequential data. Secondly, the Predictability Ceiling Hypothesis (RQ2) posits the existence of an architecture-invariant empirical upper bound on the variance explainable from end-to-end temporal history, quantifying its implications for AI-assisted data management. Empirical analysis demonstrates bounded support for RQ1: frequency-aware models achieve the highest RTT Coefficient of Determination (R2) and optimal Mean Absolute Error (MAE) at sub-100K parameters. However, their absolute predictive advantage over modern Transformers remains marginal (ΔR2 = 0.003 on RTT), with both paradigms proving indistinguishable regarding stochastic jitter. Consequently, RQ2 emerges as the primary contribution: predictive fidelity severely asymptotes at R2 ≈0.54 for propagation metrics (RTT and throughput) and saturates at R2 ≈0.31 for jitter across all sixteen structurally diverse models. This ceiling defines the maximum diagnostic gain achievable by the sequence predictors under investigation, operating exclusively on temporal 30 day LENS telemetry. We conclude that this plateau represents a fundamental informational limit of the data schema itself rather than an algorithmic deficiency, establishing that future AI-augmented edge databases should transition toward multi-modal feature fusion to breach this boundary.
AB - Effective management of Low Earth Orbit (LEO) satellite networks depends on data pipelines capable of predicting link-layer quality metrics (e.g., round-trip time (RTT), throughput (TP), and jitter etc.) at timescales suitable for real-time distributed data routing, handover, and congestion management. This study systematically benchmarks sixteen deep learning (DL) architectures representing four inductive-bias families to evaluate two formally stated hypotheses using 30 days of real Starlink telemetry comprising 417 million observations. Firstly, the Spectral Alignment Hypothesis (Research Question (RQ) 1) investigates whether architectures possessing inductive biases that explicitly decompose the quasi-periodic orbital structure of LEO dynamics systematically outperform those processing the telemetry as generic sequential data. Secondly, the Predictability Ceiling Hypothesis (RQ2) posits the existence of an architecture-invariant empirical upper bound on the variance explainable from end-to-end temporal history, quantifying its implications for AI-assisted data management. Empirical analysis demonstrates bounded support for RQ1: frequency-aware models achieve the highest RTT Coefficient of Determination (R2) and optimal Mean Absolute Error (MAE) at sub-100K parameters. However, their absolute predictive advantage over modern Transformers remains marginal (ΔR2 = 0.003 on RTT), with both paradigms proving indistinguishable regarding stochastic jitter. Consequently, RQ2 emerges as the primary contribution: predictive fidelity severely asymptotes at R2 ≈0.54 for propagation metrics (RTT and throughput) and saturates at R2 ≈0.31 for jitter across all sixteen structurally diverse models. This ceiling defines the maximum diagnostic gain achievable by the sequence predictors under investigation, operating exclusively on temporal 30 day LENS telemetry. We conclude that this plateau represents a fundamental informational limit of the data schema itself rather than an algorithmic deficiency, establishing that future AI-augmented edge databases should transition toward multi-modal feature fusion to breach this boundary.
KW - Artificial Intelligence
KW - Deep Learning
KW - LEO Satellites Telemetry
KW - Link Quality Prediction
KW - Starlink
KW - Time series Analysis
UR - https://www.scopus.com/pages/publications/105045694503
UR - https://arrow.tudublin.ie/engschelero/23/
U2 - 10.1145/3814940.3815328
DO - 10.1145/3814940.3815328
M3 - Conference contribution
AN - SCOPUS:105045694503
T3 - Proceedings of the 9th International Workshop on Artificial Intelligence Techniques for Data Management, aiDM 2026
SP - 41
EP - 53
BT - Proceedings of the 9th International Workshop on Artificial Intelligence Techniques for Data Management, aiDM 2026
PB - Association for Computing Machinery (ACM)
Y2 - 31 May 2026 through 5 June 2026
ER -