← Search

Yiye Chen

9 accepted papers

2026

Hierarchical Policy Learning via Spectral Decomposition

ICML 2026poster

In this paper, we identify a semantic decomposition in robot action sequences, separating task-level motion intent from execution-level refinements. By analyzing actions in the spectral domain using the discrete cosine transform (DCT), we observe that low-frequency components capture global motion t…

Cited by 0SourceScholar
2026

Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System

AAAI 2026technical

Scene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG2, an iterative Schema-Guided Scene-Graph reasoning framework based on multi-agent LLMs. The agents are grouped i

Cited by 0SourcePDFScholar
2025

GASP: Gaussian Avatars with Synthetic Priors

CVPR 2025poster

Gaussian Splatting has changed the game for real-time photo-realistic rendering. One of the most popular applications of Gaussian Splatting is to create animatable avatars, known as Gaussian Avatars. Recent works have pushed the boundaries of quality and rendering efficiency but suffer from two main…

Cited by 0SourcePDFScholar
2023

KGNv2: Separating Scale and Pose Prediction for Keypoint-Based 6-DoF Grasp Synthesis on RGB-D Input

IROS 2023poster

We propose an improved keypoint approach for 6-DoF grasp pose synthesis from RGB-D input. Keypoint-based grasp detection from image input demonstrated promising results in a previous study, where the visual information provided by color imagery compensates for noisy or imprecise depth measurements.…

Cited by 4SourcecodeScholar
2023

Keypoint-GraspNet: Keypoint-based 6-DoF Grasp Generation from the Monocular RGB-D input

ICRA 2023poster

The success of 6-DoF grasp learning with point cloud input is tempered by the computational costs resulting from their unordered nature and pre-processing needs for reducing the point cloud to a manageable size. These properties lead to failure on small objects with low point cloud cardinality. Inst…

Cited by 13SourcecodeScholar
2023

Planning with Sequence Models through Iterative Energy Minimization

ICLR 2023poster

Recent works have shown that language modeling can be effectively used to train reinforcement learning (RL) policies. However, the success of applying existing language models to planning, in which we wish to obtain a trajectory of actions to reach some goal, is less straightforward. The typical aut…

2023

WDiscOOD: Out-of-Distribution Detection via Whitened Linear Discriminant Analysis

ICCV 2023poster

Deep neural networks are susceptible to generating overconfident yet erroneous predictions when presented with data beyond known concepts. This challenge underscores the importance of detecting out-of-distribution (OOD) samples in the open world. In this work, we propose a novel feature-space OOD de…

Cited by 7PDFcodeScholar
2021

A Joint Network for Grasp Detection Conditioned on Natural Language Commands

ICRA 2021poster

We consider the task of grasping a target object based on a natural language command query. Previous work primarily focused on localizing the object given the query, which requires a separate grasp detection module to grasp it. The cascaded application of two pipelines incurs errors in overlapping m…

Cited by 54SourceScholar
2021

Simultaneous Multi-Level Descriptor Learning and Semantic Segmentation for Domain-Specific Relocalization

ICRA 2021poster

This paper presents a semi-supervised framework for multi-level description learning aiming for robust and accurate camera relocalization across large perception variations. Our proposed network, namely DLSSNet, simultaneously learns weakly-supervised semantic segmentation and local feature descript…

Cited by 1SourceScholar