Learnability Window in Gated Recurrent Neural Networks

2512.05790v1 cs.LG, physics.data-an 2025-12-08

Авторы:

Lorenzo Livi

Abstract

We develop a theoretical framework that explains how gating mechanisms determine the learnability window $\mathcal{H}_N$ of recurrent neural networks, defined as the largest temporal horizon over which gradient information remains statistically recoverable. While classical analyses emphasize numerical stability of Jacobian products, we show that stability alone is insufficient: learnability is governed instead by the \emph{effective learning rates} $μ_{t,\ell}$, per-lag and per-neuron quantities obtained from first-order expansions of gate-induced Jacobian products in Backpropagation Through Time. These effective learning rates act as multiplicative filters that control both the magnitude and anisotropy of gradient transport. Under heavy-tailed ($α$-stable) gradient noise, we prove that the minimal sample size required to detect a dependency at lag~$\ell$ satisfies $N(\ell)\propto f(\ell)^{-α}$, where $f(\ell)=\|μ_{t,\ell}\|_1$ is the effective learning rate envelope. This leads to an explicit formula for $\mathcal{H}_N$ and closed-form scaling laws for logarithmic, polynomial, and exponential decay of $f(\ell)$. The theory predicts that broader or more heterogeneous gate spectra produce slower decay of $f(\ell)$ and hence larger learnability windows, whereas heavier-tailed noise compresses $\mathcal{H}_N$ by slowing statistical concentration. By linking gate-induced time-scale structure, gradient noise, and sample complexity, the framework identifies the effective learning rates as the fundamental quantities that govern when -- and for how long -- gated recurrent networks can learn long-range temporal dependencies.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

Learnability Window in Gated Recurrent Neural Networks

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Temporal Graph Neural Networks for Early Anomaly Detection and Performance Predi...

GAMMA_FLOW: Guided Analysis of Multi-label spectra by MAtrix Factorization for L...

Detail Across Scales: Multi-Scale Enhancement for Full Spectrum Neural Represent...

Stochastic Clock Attention for Aligning Continuous and Ordered Sequences

OASIS: A Deep Learning Framework for Universal Spectroscopic Analysis Driven by ...

Навигация