← Search

Weiran Yao

18 accepted papers

2026

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels

ICLR 2026poster

Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its appl…

Cited by 0SourcecodeScholar
2025

APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

NeurIPS 2025poster

Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that generates verifiable and diverse multi-turn agent data. In t…

Cited by 0SourceScholar
2025

ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

EMNLP 2025

Large Action models are essential for enabling autonomous agents to perform complex tasks. However, training such models remains challenging due to the diversity of agent environments and the complexity of noisy agentic data. Existing infrastructure offers limited support for scalable, agent-specifi

2025

Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents

ICLR 2025poster

Large language model (LLM) agents have shown great potential in solving real-world software engineering (SWE) problems. The most advanced open-source SWE agent can resolve over 27% of real GitHub issues in SWE-Bench Lite. However, these sophisticated agent frameworks exhibit varying strengths, excel…

Cited by 10SourcePDFScholar
2025

PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data

ACL 2025finding

Personalization is essential for AI assistants, especially in private AI settings where models are expected to interpret users’ personal data (e.g., conversations, app usage) to understand their background, preferences, and social context. However, due to privacy concerns, existing academic research…

Cited by 23SourcePDFScholar
2025

Robust Parallel Cooperative Control of Cable-Driven Robot System via Adaptive Integral Sliding Mode

RA-L 2025

This paper develops an adaptive integral sliding mode control scheme to manipulate the cable-driven robots, realizing robust parallel cooperation performance for the variable loads. The considered cable-driven robot system (CDRS) employs multiple flexible cables to cooperatively regulate the robotic

Cited by 3SourceScholar
2025

xLAM: A Family of Large Action Models to Empower AI Agent Systems

NAACL 2025long

Autonomous agents powered by large language models (LLMs) have attracted significant research interest. However, the open-source community faces many challenges in developing specialized models for agent tasks, driven by the scarcity of high-quality agent datasets and the absence of standard protoco…

2024

APIGen: Automated PIpeline for Generating Verifiable and Diverse Function-Calling Datasets

NeurIPS 2024poster

The advancement of function-calling agent models requires diverse, reliable, and high-quality datasets. This paper presents APIGen, an automated data generation pipeline designed to synthesize high-quality datasets for function-calling applications. We leverage APIGen and collect 3,673 executable AP…

2024

CaRiNG: Learning Temporal Causal Representation under Non-Invertible Generation Process

ICML 2024poster

Identifying the underlying time-delayed latent causal processes in sequential data is vital for grasping temporal dynamics and making downstream reasoning. While some recent methods can robustly identify these latent causal variables, they rely on strict assumptions about the invertible generation p…

2024

Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization

ICLR 2024spotlight

Recent months have seen the emergence of a powerful new trend in which large language models (LLMs) are augmented to become autonomous language agents capable of performing objective oriented multi-step tasks on their own, rather than merely responding to queries from human users. Most existing lang…

2023

PLOT: Prompt Learning with Optimal Transport for Vision-Language Models

ICLR 2023top-25%

With the increasing attention to large vision-language models such as CLIP, there has been a significant amount of effort dedicated to building efficient prompts. Unlike conventional methods of only learning one single prompt, we propose to learn multiple comprehensive prompts to describe diverse ch…

2023

Temporally Disentangled Representation Learning under Unknown Nonstationarity

NeurIPS 2023poster

In unsupervised causal representation learning for sequential data with time-delayed latent causal influences, strong identifiability results for the disentanglement of causally-related latent variables have been established in stationary settings by leveraging temporal structure. However, in nonsta…

2022

Direct Trajectory Optimization of Free-Floating Space Manipulator for Reducing Spacecraft Variation

RA-L 2022

This letter investigates the direct trajectory optimization of the free-floating space manipulator (FFSM). The main purpose is to plan the joint space trajectories to reduce the spacecraft motion due to the joint rotation during the FFSM performing tasks. To improve the calculation efficiency, the a

Cited by 24SourceScholar
2022

Learning Temporally Causal Latent Processes from General Temporal Data

ICLR 2022poster

Our goal is to recover time-delayed latent causal variables and identify their relations from measured temporal data. Estimating causally-related latent variables from observations is particularly challenging as the latent variables are not uniquely recoverable in the most general case. In this work…

2022

Partial disentanglement for domain adaptation

ICML 2022spotlight

Unsupervised domain adaptation is critical to many real-world applications where label information is unavailable in the target domain. In general, without further assumptions, the joint distribution of the features and the label is not identifiable in the target domain. To address this issue, we re…

Cited by 81SourcePDFScholar
2020

Homotopic Approach for Robot Allocation Optimization Coupled With Path Constraints

RA-L 2020

This letter investigates a special task allocation problem with constraints of path planning. The path planning process involved is defined as the follow-up step of task allocation. The allocation problem is coupled with path constraints, which will affect the utility of allocation solution. A homot

Cited by 8SourceScholar
2019

Automated Aortic Pressure Regulation in ex vivo Heart Perfusion

ICRA 2019poster

This paper presents the first system for automated ex vivo perfusion of an isolated heart and regulating the heart's aortic pressure (AoP). An adaptive controller was developed for AoP regulation and maintained the heart's physiological aerobic metabolism. A mathematical model of the perfusion syste…

Cited by 0SourceScholar