Near Optimal Inference for the Best-Performing Algorithm

2508.05173v1 cs.LG, stat.ML 2025-08-09

Авторы:

Amichai Painsky

Резюме на русском

Набор конкурирующих машинно обучаемых алгоритмов может быть оценен по результатам на наборе данных. Цель — определить, какой алгоритм будет скорее всего дать наилучший результат на будущих неизвестных данных. Обычно выбирается тот алгоритм, который показал лучшую результативность на текущем наборе. Однако в некоторых случаях разница в показателях между алгоритмами мала, и некоторые могут быть включены в рассмотрение. Этот вопрос сформулирован как задача выбора подмножества для многочленных распределений. Задача: найти минимальное подмножество символов, включающее наиболее частотный символ в общей популяции с достаточно высоким уровнем уверенности. Работа предлагает новую модель для решения этой задачи. Она включает асимптотические и локальные схемы, превышающие существующие методы по точности и эффективности. Также доказаны соответствующие нижние оценки, подтверждающие выгодность предложенных подходов.

Abstract

Consider a collection of competing machine learning algorithms. Given their performance on a benchmark of datasets, we would like to identify the best performing algorithm. Specifically, which algorithm is most likely to rank highest on a future, unseen dataset. A natural approach is to select the algorithm that demonstrates the best performance on the benchmark. However, in many cases the performance differences are marginal and additional candidates may also be considered. This problem is formulated as subset selection for multinomial distributions. Formally, given a sample from a countable alphabet, our goal is to identify a minimal subset of symbols that includes the most frequent symbol in the population with high confidence. In this work, we introduce a novel framework for the subset selection problem. We provide both asymptotic and finite-sample schemes that significantly improve upon currently known methods. In addition, we provide matching lower bounds, demonstrating the favorable performance of our proposed schemes.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

Near Optimal Inference for the Best-Performing Algorithm

Авторы:

Резюме на русском

Abstract

Ссылки и действия

Связанные статьи

Breaking Determinism: Stochastic Modeling for Reliable Off-Policy Evaluation in ...

Tuning-Free Structured Sparse Recovery of Multiple Measurement Vectors using Imp...

GaussDetect-LiNGAM:Causal Direction Identification without Gaussianity test

Parameter-Efficient Augment Plugin for Class-Incremental Learning

Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sampl...

Навигация