A Complete Characterization of Learnability for Stochastic Noisy Bandits

AmazUtah_NLP at SemEval-2024 Task 9: A MultiChoice Question Answering System for Commonsense Defying Reasoning


View a PDF of the paper titled A Complete Characterization of Learnability for Stochastic Noisy Bandits, by Steve Hanneke and 1 other authors

View PDF

Abstract:We study the stochastic noisy bandit problem with an unknown reward function $f^*$ in a known function class $mathcal{F}$. Formally, a model $M$ maps arms $pi$ to a probability distribution $M(pi)$ of reward. A model class $mathcal{M}$ is a collection of models. For each model $M$, define its mean reward function $f^M(pi)=mathbb{E}_{r sim M(pi)}[r]$. In the bandit learning problem, we proceed in rounds, pulling one arm $pi$ each round and observing a reward sampled from $M(pi)$. With knowledge of $mathcal{M}$, supposing that the true model $Min mathcal{M}$, the objective is to identify an arm $hat{pi}$ of near-maximal mean reward $f^M(hat{pi})$ with high probability in a bounded number of rounds. If this is possible, then the model class is said to be learnable.

Importantly, a result of cite{hanneke2023bandit} shows there exist model classes for which learnability is undecidable. However, the model class they consider features deterministic rewards, and they raise the question of whether learnability is decidable for classes containing sufficiently noisy models. For the first time, we answer this question in the positive by giving a complete characterization of learnability for model classes with arbitrary noise. In addition to that, we also describe the full spectrum of possible optimal query complexities. Further, we prove adaptivity is sometimes necessary to achieve the optimal query complexity. Last, we revisit an important complexity measure for interactive decision making, the Decision-Estimation-Coefficient citep{foster2021statistical,foster2023tight}, and propose a new variant of the DEC which also characterizes learnability in this setting.

Submission history

From: Kun Wang [view email]
[v1]
Sat, 12 Oct 2024 17:23:34 UTC (41 KB)
[v2]
Fri, 17 Jan 2025 00:25:18 UTC (58 KB)



Source link
lol

By stp2y

Leave a Reply

Your email address will not be published. Required fields are marked *

No widgets found. Go to Widget page and add the widget in Offcanvas Sidebar Widget Area.