2021
Joint Online Learning and Decision-making via Dual Mirror Descent
ICML 2021spotlight
We consider an online revenue maximization problem over a finite time horizon subject to lower and upper bounds on cost. At each period, an agent receives a context vector sampled i.i.d. from an unknown distribution and needs to make a decision adaptively. The revenue and cost functions depend on th…