Learnability Window in Gated Recurrent Neural Networks

2512.05790v1 cs.LG, physics.data-an 2025-12-08
Авторы:

Lorenzo Livi

Abstract

We develop a theoretical framework that explains how gating mechanisms determine the learnability window $\mathcal{H}_N$ of recurrent neural networks, defined as the largest temporal horizon over which gradient information remains statistically recoverable. While classical analyses emphasize numerical stability of Jacobian products, we show that stability alone is insufficient: learnability is governed instead by the \emph{effective learning rates} $μ_{t,\ell}$, per-lag and per-neuron quantities obtained from first-order expansions of gate-induced Jacobian products in Backpropagation Through Time. These effective learning rates act as multiplicative filters that control both the magnitude and anisotropy of gradient transport. Under heavy-tailed ($α$-stable) gradient noise, we prove that the minimal sample size required to detect a dependency at lag~$\ell$ satisfies $N(\ell)\propto f(\ell)^{-α}$, where $f(\ell)=\|μ_{t,\ell}\|_1$ is the effective learning rate envelope. This leads to an explicit formula for $\mathcal{H}_N$ and closed-form scaling laws for logarithmic, polynomial, and exponential decay of $f(\ell)$. The theory predicts that broader or more heterogeneous gate spectra produce slower decay of $f(\ell)$ and hence larger learnability windows, whereas heavier-tailed noise compresses $\mathcal{H}_N$ by slowing statistical concentration. By linking gate-induced time-scale structure, gradient noise, and sample complexity, the framework identifies the effective learning rates as the fundamental quantities that govern when -- and for how long -- gated recurrent networks can learn long-range temporal dependencies.

Ссылки и действия

Связанные статьи

Detail Across Scales: Multi-Scale Enhancement for Full Spectrum Neural Represent...

## Контекст Implicit neural representations (INRs) представляют собой мощный подход к кодированию данных, использующий н...

2025-09-23

Stochastic Clock Attention for Aligning Continuous and Ordered Sequences

## Контекст Современные подходы в обработке и анализе данных часто сталкиваются с задачами построения моделей, которые о...

2025-09-20

OASIS: A Deep Learning Framework for Universal Spectroscopic Analysis Driven by ...

## Контекст Спектроскопические данные широко распространены в различных научных и инженерных областях, требуя эффективн...

2025-09-17