← Search

Qing Da

2 accepted papers

2026

Reducing Belief Deviation in Reinforcement Learning for Active Reasoning

ICLR 2026oral

Active reasoning requires large language models (LLMs) to interact with external sources and strategically gather information to solve problems. Central to this process is belief tracking: maintaining a coherent understanding of the problem state and the missing information toward the solution. Howe…

Cited by 0SourcecodeScholar
2021

A Primal-Dual Online Algorithm for Online Matching Problem in Dynamic Environments

AAAI 2021technical

Recently, the online matching problem has attracted much attention due to its wide application on real-world decision-making scenarios. In stationary environments, by adopting the stochastic user arrival model, existing methods are proposed to learn dual optimal prices and are shown to achieve a fas…

Cited by 2SourcePDFScholar