← Search

Shaoyuan Xu

6 accepted papers

2026

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems

ICML 2026poster

Multi-agent systems (MAS) are increasingly capable of tackling complex real-world tasks, yet their reliance on inter-agent coordination, tool use, and long-horizon reasoning makes error recognition particularly challenging. Minor errors can propagate across agents, escalating into task failures whil…

Cited by 0SourceScholar
2026

CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization

CVPR 2026

Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful visual reasoning: models may invoke tools on irrelevant regions or ignore tool outputs entirely, yet still guess the cor

Cited by 0SourcecodeScholar
2025

STIMULUS: Achieving Fast Convergence and Low Sample Complexity in Stochastic Multi-Objective Learning

UAI 2025

Recently, multi-objective optimization (MOO) has gained attention for its broad applications in ML, operations research, and engineering. However, MOO algorithm design remains in its infancy and many existing MOO methods suffer from unsatisfactory convergence rate and sample complexity performance.

Cited by 0SourcePDFScholar
2024

Q-Tuning: Queue-based Prompt Tuning for Lifelong Few-shot Language Learning

NAACL 2024findings

This paper introduces Q-tuning, a novel approach for continual prompt tuning that enables the lifelong learning of a pre-trained language model. When learning a new task, Q-tuning trains a task-specific prompt by adding it to a prompt queue consisting of the prompts from older tasks. To better trans…

Cited by 5SourcePDFScholar
2023

KG-FLIP: Knowledge-guided Fashion-domain Language-Image Pre-training for E-commerce

ACL 2023industry

Various Vision-Language Pre-training (VLP) models (e.g., CLIP, BLIP) have sprung up and dramatically advanced the benchmarks for public general-domain datasets (e.g., COCO, Flickr30k). Such models usually learn the cross-modal alignment from large-scale well-aligned image-text datasets without lever…

Cited by 11SourcePDFScholar