← Search

Yifei Yang

25 accepted papers

2026

Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I)

ICRA 2026poster

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation…

Cited by 0codeScholar
2026

Toward Embodiment Equivariant Vision-Language-Action Policy

ICRA 2026poster

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying le…

2025

DORec: Decomposed Object Reconstruction and Segmentation Utilizing 2D Self-Supervised Features

RA-L 2025

Recovering 3D geometry and textures of individual objects is crucial for many robotics applications, such as manipulation, pose estimation, and autonomous driving. However, decomposing a target object from a complex background is challenging. Most existing approaches rely on costly manual labels to

Cited by 1SourceScholar
2025

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

IROS 2025

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data

Cited by 2SourceScholar
2025

FreDF: Learning to Forecast in the Frequency Domain

ICLR 2025poster

Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting mode…

2025

PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization

ACL 2025finding

Large Language Model (LLM) agents have demonstrated impressive capabilities in handling complex interactive problems. Existing LLM agents mainly generate natural language plans to guide reasoning, which is verbose and inefficient. NL plans are also tailored to specific tasks and restrict agents’ abi…

2025

SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering

AAAI 2025technical

Safety alignment is indispensable for Large language models (LLMs) to defend threats from malicious instructions. However, recent researches reveal safety-aligned LLMs tend to reject benign queries due to the exaggerated safety issue, limiting their helpfulness. In this paper, we propose a Safety-Co…

2024

CMMLU: Measuring massive multitask language understanding in Chinese

ACL 2024findings

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance is becoming more important and more challenging. This paper aims to address this issue for Mandarin Chinese in the form of CMMLU, a comprehensive Chinese benchmark that covers various subjects, incl…

2024

Drag Your Noise: Interactive Point-based Editing via Diffusion Semantic Propagation

CVPR 2024poster

Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work DragDiffusion updates the diffusion latent map in response to user inputs causing global latent map alterations. This results in imprecise preservation of th…

2024

Improving Hyperbolic Representations via Gromov-Wasserstein Regularization

ECCV 2024poster

"Hyperbolic representations have shown remarkable efficacy in modeling inherent hierarchies and complexities within data structures. Hyperbolic neural networks have been commonly applied for learning such representations from data, but they often fall short in preserving the geometric structures of…

2024

ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene Reconstruction

IROS 2024poster

The joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrizat…

Cited by 0SourceScholar
2023

Learning Transformation-Predictive Representations for Detection and Description of Local Features

CVPR 2023poster

The task of key-points detection and description is to estimate the stable location and discriminative representation of local features, which is essential for image matching. However, either the rough hard positive or negative labels generated from one-to-one correspondences among images bring indi…

Cited by 11SourcePDFScholar
2023

Open-Set Object Detection Using Classification-Free Object Proposal and Instance-Level Contrastive Learning

RA-L 2023

Detecting both known and unknown objects is a fundamental skill for robot manipulation in unstructured environments. Open-set object detection (OSOD) is a promising direction to handle the problem consisting of two subtasks: objects and background separation, and open-set object classification. In t

Cited by 21SourceScholar
2023

RefGPT: Dialogue Generation of GPT, by GPT, and for GPT

EMNLP 2023long findings

Large Language Models (LLMs) have attained the impressive capability to resolve a wide range of NLP tasks by fine-tuning high-quality instruction data. However, collecting human-written data of high quality, especially multi-turn dialogues, is expensive and unattainable for most people. Though previ…

Cited by 0SourcecodeScholar
2023

UrbanGIRAFFE: Representing Urban Scenes as Compositional Generative Neural Feature Fields

ICCV 2023poster

Generating photorealistic images with controllable camera pose and scene contents is essential for many applications including AR/VR and simulation. Despite the fact that rapid progress has been made in 3D-aware generative models, most existing methods focus on object-centric images and are not appl…

Cited by 17PDFScholar
2022

Rethinking Controllable Variational Autoencoders

CVPR 2022poster

The Controllable Variational Autoencoder (ControlVAE) combines automatic control theory with the basic VAE model to manipulate the KL-divergence for overcoming posterior collapse and learning disentangled representations. It has shown success in a variety of applications, such as image generation, d…

Cited by 14PDFScholar
2019

Vehicle Re-Identification in Aerial Imagery: Dataset and Approach

ICCV 2019poster

In this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured…

Cited by 78PDFScholar
2018

Flexible Multi-Group Single-Carrier Modulation: Optimal Subcarrier Grouping and Rate Maximization

ICASSP 2018accepted

Orthogonal frequency division multiplexing (OFDM) and single-carrier frequency domain equalization (SC-FDE) are two commonly adopted modulation schemes for frequency-selective channels. Compared to SC-FDE, OFDM generally achieves higher data rate, but at the cost of higher transmit signal peak-to-av…

Cited by 0SourceScholar
2018

Low Complexity Implementation of Carrier and Symbol Timing Synchronization for a Fully Digital Downhole Telemetry System

ICASSP 2018accepted

This paper presents a low complexity implementation of carrier phase and symbol timing synchronization on an OMAP-LI37 DSP-ARM dual core processor for a M -ary phase-shift-keying (M -PSK) based fully digital downhole telemetry system. The synchronization uses a data-aided algorithm (DA) that exploit…

Cited by 0SourceScholar