← Search

Simon Yu

8 accepted papers

2026

One Skill, Many Websites: Learning Generalizable Skills Through Polymorphic Abstraction

ICLR 2026poster

Large language models (LLMs) are moving beyond static uses and are now powering agents that learn during their interaction with external environments. For example, agents can learn reusable skills while navigating web pages or toggling new tools. However, existing methods for skill learning often cr…

Cited by 0SourcecodeScholar
2026

SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning

ICLR 2026poster

Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approaches depend on human-curated problem-answer pairs and domain-specific reward engineering. We introduce SPIRAL, a self-play…

Cited by 0SourcecodeScholar
2026

Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

ICML 2026poster

LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in multi-turn interactions and employ diverse tools, introducing new risks overlooked by existing benchmarks. To systematica…

Cited by 0SourceScholar
2026

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity

ICML 2026poster

Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level driver: typicality bias in preference data, whereby annotators systematically…

Cited by 0SourceScholar
2024

Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?

EMNLP 2024main

Multilingual large language models are designed, claimed, and expected to cater to speakers of varied languages. We hypothesise that the current practices of fine-tuning and evaluating these models may not perfectly align with this objective owing to a heavy reliance on translation, which cannot cov…

2019

Trajectory Estimation for Geo-Fencing Applications on Small-Size Fixed-Wing UAVs

IROS 2019poster

The steadily increasing popularity of Unmanned Aerial Vehicles (UAVs) is creating new opportunities in diverse fields of technology and business. However, this increase of popularity also raises safety concerns. To tackle the primary concern of keeping the UAV inside a designated region, a novel tra…

Cited by 4SourceScholar