← Search

Hao Qiu

11 accepted papers

2026

Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory

ICML 2026poster

In this paper, we study dynamic regret in unconstrained online convex optimization (OCO) with movement costs. Specifically, we generalize the standard setting by allowing the movement cost coefficients $\lambda_t$ to vary arbitrarily over time. Our main contribution is a novel algorithm that establi…

Cited by 0SourceScholar
2025

EgoBlind: Towards Egocentric Visual Assistance for the Blind

NeurIPS 2025poster

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It…

Cited by 0SourcecodeScholar
2025

Exploiting Curvature in Online Convex Optimization with Delayed Feedback

ICML 2025poster

In this work, we study the online convex optimization problem with curved losses and delayed feedback. When losses are strongly convex, existing approaches obtain regret bounds of order $d_{\max} \ln T$, where $d_{\max}$ is the maximum delay and $T$ is the time horizon. However, in many cases, this…

Cited by 0SourcePDFScholar
2025

Federated Multi-armed Bandits with Efficient Bit-Level Communications

NeurIPS 2025poster

In this work, we study the federated multi-armed bandit (FMAB) problem, where a set of distributed agents collaboratively aim to minimize cumulative regret while interacting with a shared set of arms. Unlike traditional centralized bandit models, agents in FMAB settings are connected via a communica…

Cited by 0SourceScholar
2025

Near-Optimal Regret Bounds for Federated Multi-armed Bandits with Fully Distributed Communication

UAI 2025

In this paper, we focus on the research of federated multi-armed bandit (FMAB) problems where agents can only communicate with their neighbors. All agents aim to solve a common multi-armed bandit (MAB) problem to minimize individual regrets, while group regret can also be minimized. In a federated b

Cited by 0SourcePDFScholar
2023

Delayed Bandits: When Do Intermediate Observations Help?

ICML 2023poster

We study a $K$-armed bandit with delayed feedback and intermediate observations. We consider a model, where intermediate observations have a form of a finite state, which is observed immediately after taking an action, whereas the loss is observed after an adversarially chosen delay. We show that th…

Cited by 2SourcePDFScholar
2023

Trading-Off Payments and Accuracy in Online Classification with Paid Stochastic Experts

ICML 2023poster

We investigate online classification with paid stochastic experts. Here, before making their prediction, each expert must be paid. The amount that we pay each expert directly influences the accuracy of their prediction through some unknown Lipschitz ``productivity'' function. In each round, the lear…

Cited by 0SourcePDFScholar
2020

An Untethered 216-mg Insect-Sized Jumping Robot with Wireless Power Transmission

IROS 2020poster

We present the first demonstration of a battery-free untethered wirelessly powered sub-gram jumping robot on an insect-scale. In order to operate the insect-sized robot autonomously, the limitation in battery use emphasizes the need for a wireless power transmission system as an onboard power soluti…

Cited by 19SourceScholar