Expand ↗
Page list (1404)

Online Learning

A sequential decision model: at each round a learner picks an action, then suffers a loss chosen possibly adversarially, and updates — performance measured by regret against the best fixed action in hindsight. The setting in which no-regret dynamics (e.g. FTRL) are analysed, and the lens through which “Do LLM Agents Have Regret?” asks whether language-model agents behave like rational repeated-game players.

In this vault

Backlinks