IJCAI 20260 citations

Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach

Michael Ledford, William Regli

Abstract

In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is fully censored. This censorship creates an identifiability problem: agents cannot distinguish stochastic failure from insufficient coordination. We formalize this setting as the Threshold-Activated Cooperative Multi-Armed Bandit (TAC-MAB) and analyze it under both centralized and decentralized coordination. We show that a centralized algorithm (C-TAC) achieves cumulative regret O(log T), decomposed into a structural-search term that captures the cost of resolving feasibility under censored feedback and a statistical-monitoring term for value estimation. We then introduce D-TAC, a decentralized event-triggered protocol in which agents synchronize only when their structural beliefs change. Empirically, D-TAC achieves a 23x reduction in communication relative to the centralized baseline while preserving feasibility alignment under conservative belief fusion. These results characterize the coordination cost of learning under censored feedback and show that near-centralized communication efficiency is achievable without continuous synchronization.

Machine Learning: Multi-armed banditsAgent-based and Multi-agent Systems: Coordination and cooperationAgent-based and Multi-agent Systems: Multi-agent learningMachine Learning: Online learningAgent-based and Multi-agent Systems: Agent communication
BibTeX
@inproceedings{ijcai2026_costofstructural,
  title = {Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach},
  author = {Michael Ledford and William Regli},
  booktitle = {IJCAI 2026},
  year = {2026}
}
Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach · IJCAI 2026