← Search

Yushi Yang

11 accepted papers

2026

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

ICML 2026poster

Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, however, makes them vulnerable to prompt injection attacks: adversarial instructions hidden in interface elements that persuad…

Cited by 0SourceScholar
2026

LeHome: A Simulation Environment for Deformable Object Manipulation in Household Scenarios

ICRA 2026poster

Household environments present one of the most common, impactful yet challenging application domains for robotics. Within household scenarios, manipulating deformable objects is particularly difficult, both in simulation and real-world execution, due to varied categories and shapes, complex dynamics…

2025

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

EMNLP 2025

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations—attributing its effects solely to dampened toxic neurons in the MLP layers—are incomplete

Cited by 0SourcePDFScholar
2025

LLMs Don’t Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

EMNLP 2025

To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have p

2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar
2025

Unidirectional Point-Voxel Fusion for Enhanced 3D Single Object Tracking

IROS 2025

Sparse point-based trackers struggle with texture-less and incomplete point clouds. Conversely, dense voxel-based trackers have richer spatial and semantic information, but filtering out interference from complex backgrounds remains a challenge. Additionally, there is still a gap between point and v

Cited by 0SourceScholar
2024

Enhancing 3D Single Object Tracking with Efficient Point Cloud Segmentation

IROS 2024poster

3D single object tracking (SOT) based on point cloud has attracted much attention due to its important role in machine vision and autonomous driving. Recently, M2-Track proposes a two-stage tracking structure centered on motion, but they ignore the effect of segmentation errors in sparse point cloud…

Cited by 0SourceScholar
2024

Integrating Scaling Strategy and Central Guided Voting for 3D Point Cloud Object Tracking

RA-L 2024

LiDAR-based 3D single object tracking has received remarkable attention due to its crucial role in robotics and autonomous driving. Most of them are based on hierarchical feature structures from PointNet++. However, existing based-stratified structure trackers ignore the fact that non-linearities in

Cited by 4SourceScholar