← Search

Yuchen Yang

30 accepted papers

2026

AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models

ICML 2026poster

Reinforcement learning (RL) based post-training for large language models (LLMs) is computationally expensive, as it generates many rollout sequences that could frequently share long token prefixes. Existing RL frameworks usually process these sequences independently, repeatedly recomputing identica…

Cited by 0SourceScholar
2026

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

ICML 2026poster

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind, remaining fragmented and narrowly focused. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for Compo…

Cited by 0SourceScholar
2026

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

ICLR 2026poster

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to specific downstream tasks, typically involves further training on corresponding datasets. However, a fundamental discrepancy…

Cited by 0SourceScholar
2026

From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

CVPR 2026

End-to-end multi-object tracking (MOT) methods have recently achieved remarkable progress by unifying detection and association within a single framework. Despite their strong detection performance, these methods suffer from relatively low association accuracy. Through detailed analysis, we observe

Cited by 0SourcecodeScholar
2026

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

ICML 2026oral

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question–answer (QA) pairs from a limited number of manually annotated datasets, rather than sy…

Cited by 0SourceScholar
2026

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

AAAI 2026technical

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research

Cited by 0SourcePDFScholar
2026

Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

ICML 2026poster

Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely on empirical evidence and lack a rigorous theoretical understanding of the condi…

Cited by 0SourceScholar
2026

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling

ICML 2026poster

While Mixture-of-Experts (MoE) architectures substantially bolster the expressive power of large-language models, their prohibitive memory footprint severely impedes the practical deployment on resource-constrained edge devices, especially when model behavior must be preserved without relying on los…

Cited by 0SourceScholar
2025

CL-DiffPhyCon: Closed-loop Diffusion Control of Complex Physical Systems

ICLR 2025poster

The control problems of complex physical systems have broad applications in science and engineering. Previous studies have shown that generative control methods based on diffusion models offer significant advantages for solving these problems. However, existing generative control approaches face ch…

2025

DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models

EMNLP 2025

The reliability of large language models remains a critical challenge, particularly due to their susceptibility to hallucinations and factual inaccuracies during text generation. Existing solutions either underutilize models’ self-correction with preemptive strategies or use costly post-hoc verifica

Cited by 0SourcePDFScholar
2025

GMCL: Graph-Enhanced Multimodal Contrastive Learning for Rumor Detection

ICASSP 2025accepted

Multimedia rumor content has been widely disseminated with the rise of generative technologies. Existing rumor detection approaches typically focus independently on multi-modal data (such as text and images) or social structure analysis, and only a few researchers have attempted to integrate all thr…

Cited by 0SourceScholar
2025

Hierarchical Trajectory Planning Method for Piano-Playing Robot

IROS 2025

Piano-playing tasks, which effectively demonstrate bimanual coordination capabilities in humanoid robots, are increasingly becoming a research focus. However, prior research has predominantly focused on Cartesian space trajectory planning without adequately addressing real-world obstacle avoidance c

Cited by 0SourceScholar
2025

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

IJCAI 2025

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and larg

2024

A Combination of a Controllable Clutch and an Oscillating Slider Crank Mechanism for Ease of Direct-Teaching with Various Payloads

ICRA 2024poster

Direct teaching is a straightforward way of teaching new motion to robots. Active methods with torque sensors, for example, can be used so that the robot can follow the movements of the human, but such methods introduce delays. Alternatively, series clutch actuators are easily backdrivable without d…

Cited by 0SourceScholar
2024

CF-TCIR: A Compositor-Free Framework for Hierarchical Text-Conditioned Image Retrieval

ACL 2024findings

In text-conditioned image retrieval (TCIR), the combination of a reference image and modification text forms a query tuple, aiming to locate the most congruent target image within a dataset. The advantages of rich image semantic information and text flexibility are combined in this manner for more a…

Cited by 1SourcePDFScholar
2024

CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression Segmentation

ICRA 2024poster

The classical human-robot interface in uncalibrated image-based visual servoing (UIBVS) relies on either human annotations or semantic segmentation with categorical labels. Both methods fail to match natural human communication and convey rich semantics in manipulation tasks as effectively as natura…

Cited by 1SourceScholar
2024

DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics

ACL 2024findings

Structured data offers an efficient means of organizing information. Exsisting text-serialization based methods for processing structured data using large language models (LLMs) are not designed to explicitly capture the heterogeneity of structured data. Such methods are suboptimal for LLMs to proce…

Cited by 1SourcePDFScholar
2024

HSDreport: Heart Sound Diagnosis with Echocardiography Reports

EMNLP 2024finding

Heart sound auscultation holds significant importance in the diagnosis of congenital heart disease. However, existing methods for Heart Sound Diagnosis (HSD) tasks are predominantly limited to a few fixed categories, framing the HSD task as a rigid classification problem that does not fully align wi…

Cited by 0SourcePDFScholar
2024

Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation

ECCV 2024poster

"Automatic estimation of 3D human pose from monocular RGB images is a challenging and unsolved problem in computer vision. In a supervised manner, approaches heavily rely on laborious annotations and present hampered generalization ability due to the limited diversity of 3D pose datasets. To address…

2024

Monocular Localization with Semantics Map for Autonomous Vehicles

ICRA 2024poster

Accurate and robust localization remains a significant challenge for autonomous vehicles. The cost of sensors and limitations in local computational efficiency make it difficult to scale to large commercial applications. Traditional vision-based approaches focus on texture features that are suscepti…

Cited by 0SourceScholar
2024

Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation

ICML 2024poster

Fine-tuning pretrained large models to downstream tasks is an important problem, which however suffers from huge memory overhead due to large-scale parameters. This work strives to reduce memory overhead in fine-tuning from perspectives of activation function and layer normalization. To this end, we…

2024

Robust Noisy Correspondence Learning with Equivariant Similarity Consistency

CVPR 2024poster

The surge in multi-modal data has propelled cross-modal matching to the forefront of research interest. However the challenge lies in the laborious and expensive process of curating a large and accurately matched multimodal dataset. Commonly sourced from the Internet these datasets often suffer from…

Cited by 6SourcePDFScholar
2024

Robustness-Guided Image Synthesis for Data-Free Quantization

AAAI 2024technical

Quantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the…

Cited by 4SourcePDFScholar
2023

DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds

ICCV 2023poster

Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of offboard 3D detectors is not explored mainly due to two reasons: (1) the onboard multi-object tracker cannot generate sufficient com…

Cited by 35PDFcodeScholar
2023

LoGoNet: Towards Accurate 3D Object Detection With Local-to-Global Cross-Modal Fusion

CVPR 2023poster

LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding…

2023

UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase

ICCV 2023poster

Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more rob…

Cited by 46PDFcodeScholar
2022

Addressing Heterogeneity in Federated Learning via Distributional Transformation

ECCV 2022poster

"Federated learning (FL) allows multiple clients to collaboratively train a deep learning model. One major challenge of FL is when data distribution is heterogeneous, i.e., differs from one client to another. Existing personalized FL algorithms are only applicable to narrow cases, e.g., one or two d…

2022

Pose Refinement with Joint Optimization of Visual Points and Lines

IROS 2022poster

High-precision camera re-localization technology in a pre-established 3D environment map is the basis for many tasks, such as Augmented Reality, Robotics and Autonomous Driving. The point-based visual re-localization approaches are well-developed in recent decades, but are insufficient in some featu…

Cited by 22SourceScholar
2021

Retrieval and Localization with Observation Constraints

ICRA 2021poster

Accurate visual re-localization is very critical to many artificial intelligence applications, such as augmented reality, virtual reality, robotics and autonomous driving. To accomplish this task, we propose an integrated visual re-localization method called RLOCS by combining image retrieval, seman…

Cited by 10SourceScholar