← Search

Xiaoqing Wang

3 accepted papers

2026

IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning

ICML 2026poster

Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We revisit token-efficient post-training and argue that existing sequence-level reward-shaping methods offer limited control over how reasoning effort is …

Cited by 0SourceScholar
2026

Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems

AAAI 2026technical

The rapid advancement of Large Language Model (LLM)-driven multi-agent systems has significantly streamlined software developing tasks, enabling users with little technical expertise to develop executable applications. While these systems democratize software creation through natural language requir

Cited by 0SourcePDFScholar
2023

Generalize Learned Heuristics to Solve Large-scale Vehicle Routing Problems in Real-time

ICLR 2023poster

Large-scale Vehicle Routing Problems (VRPs) are widely used in logistics, transportation, supply chain, and robotic systems. Recently, data-driven VRP heuristics are proposed to generate real-time VRP solutions with up to 100 nodes. Despite this progress, current heuristics for large-scale VRPs stil…

Cited by 58SourcePDFScholar