Unsupervised lexicon learning from speech is limited by representations rather than clustering
2510.09225v1
eess.AS, cs.CL, cs.SD
2025-10-14
Авторы:
Danel Adendorff, Simon Malan, Herman Kamper
Abstract
Zero-resource word segmentation and clustering systems aim to tokenise speech
into word-like units without access to text labels. Despite progress, the
induced lexicons are still far from perfect. In an idealised setting with gold
word boundaries, we ask whether performance is limited by the representation of
word segments, or by the clustering methods that group them into word-like
types. We combine a range of self-supervised speech features
(continuous/discrete, frame/word-level) with different clustering methods
(K-means, hierarchical, graph-based) on English and Mandarin data. The best
system uses graph clustering with dynamic time warping on continuous features.
Faster alternatives use graph clustering with cosine distance on averaged
continuous features or edit distance on discrete unit sequences. Through
controlled experiments that isolate either the representations or the
clustering method, we demonstrate that representation variability across
segments of the same word type -- rather than clustering -- is the primary
factor limiting performance.
Ссылки и действия
Дополнительные ресурсы: