Convergence for Discrete Parameter Updates

2512.04051v1 cs.LG, math.OC 2025-12-05

Авторы:

Paul Wilson, Fabio Zanasi, George Constantinides

Abstract

Modern deep learning models require immense computational resources, motivating research into low-precision training. Quantised training addresses this by representing training components in low-bit integers, but typically relies on discretising real-valued updates. We introduce an alternative approach where the update rule itself is discrete, avoiding the quantisation of continuous updates by design. We establish convergence guarantees for a general class of such discrete schemes, and present a multinomial update rule as a concrete example, supported by empirical evaluation. This perspective opens new avenues for efficient training, particularly for models with inherently discrete structure.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

Найти цитирования в Google Scholar
Поиск в Semantic Scholar
Другие статьи категории cs.LG, math.OC

Convergence for Discrete Parameter Updates

Авторы:

Abstract

Ссылки и действия

Связанные статьи

The Geometry of Intelligence: Deterministic Functional Topology as a Foundation ...

Beyond Scaffold: A Unified Spatio-Temporal Gradient Tracking Method

Risk-Sensitive Q-Learning in Continuous Time with Application to Dynamic Portfol...

ARM-Explainer -- Explaining and improving graph neural network predictions for t...

Mean-Field Limits for Two-Layer Neural Networks Trained with Consensus-Based Opt...

Навигация