PAPER DIGEST
Most Influential AISTATS 2011 Paper · 2026-03 edition

Contextual Bandits With Linear Payoff Functions

Wei Chu; Lihong Li; Lev Reyzin; Robert Schapire

Venue
Conference on Artificial Intelligence and Statistics (AISTATS) 2011
Recognition
Most Influential AISTATS 2011 Paper (Rank No. 4)
Edition
2026-03
Impact factor
9
Certificate ID
2b274f9186cf66a8

Abstract

In this paper we study the contextual bandit problem (also known as the multi-armed bandit problem with expert advice) for linear payoff functions. For T rounds, K actions, and d dimensional feature vectors, we prove an O(\sqrtTd\ln^3(KT\ln(T)/δ)) regret bound that holds with probability 1-δfor the simplest known (both conceptually and computationally) efficient upper confidence bound algorithm for this problem. We also prove a lower bound of Ω(\sqrtTd) for this setting, matching the upper bound up to logarithmic factors. [pdf]

Download PDF certificate