← Search

Yi-Ting Chen

36 accepted papers

2026

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

AAAI 2026technical

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances, resulting in frequent manipulation failures. We propose Afforda

Cited by 0SourcePDFScholar
2026

Controllable Collision Scenario Generation Via Collision Pattern Prediction

ICRA 2026poster

Evaluating the safety of autonomous vehicles (AVs) requires diverse, safety-critical scenarios, with collisions being especially important yet rare and unsafe to collect in the real world. Therefore, the community has been focusing on generating safety-critical scenarios in simulation. However, cont…

2026

GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks

ICRA 2026poster

Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especially learn-from-demonstration methods, still struggle to handle diverse and dynamic food states, which often results in spillage and reduced reliability.…

2026

HetroD: A High-Fidelity Drone Dataset and Benchmark for Autonomous Driving in Heterogeneous Traffic

ICRA 2026poster

We present HetroD, a dataset and benchmark for developing autonomous driving systems in heterogeneous environments. HetroD targets the critical challenge of navigating real-world heterogeneous traffic dominated by vulnerable road users (VRUs), including pedestrians, cyclists, motorcyclists, and vehi…

2026

Uncertainty-Aware Vision-Based Risk Object Identification Via Conformal Risk Tube Prediction

ICRA 2026poster

We study object importance-based vision risk object identification (Vision-ROI), a key capability for intelligent driving systems. Existing approaches are deterministic and ignore uncertainty, potentially compromising safety. For example, fixed decision thresholds in ambiguous scenarios can cause pr…

2025

ATARS: An Aerial Traffic Atomic Activity Recognition and Temporal Segmentation Dataset

IROS 2025

Traffic Atomic Activity, which describes traffic patterns for topological intersection dynamics, is a crucial topic for the advancement of intelligent driving systems. However, existing atomic activity datasets are collected from an egocentric view, which cannot support the scenarios where traffic a

Cited by 1SourcecodeScholar
2025

ArticuBot: Learning Universal Articulated Object Manipulation Policy via Large Scale Simulation

RSS 2025poster

This paper presents ArticuBot, in which a single learned policy enables a robotics system to open diverse categories of unseen articulated objects in the real world. This task has long been challenging for robotics due to the large variations in the geometry, size, and articulation types of such ob…

Cited by 0PDFScholar
2025

Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework

ICCV 2025poster

We propose 3D Super Resolution (3DSR), a novel 3D Gaussian-splatting-based super-resolution framework that leverages off-the-shelf diffusion-based 2D super-resolution models. 3DSR encourages 3D consistency across views via the use of an explicit 3D Gaussian-splatting-based scene representation. This…

Cited by 0SourcePDFScholar
2025

Potential Fields as Scene Affordance for Behavior Change-Based Visual Risk Object Identification

ICRA 2025

We study behavior change-based visual risk object identification (Visual-ROI), a critical framework designed to detect potential hazards for intelligent driving systems. Existing methods often show significant limitations in spatial accuracy and temporal consistency, stemming from an incomplete unde

Cited by 3SourceScholar
2025

RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration Network

CVPR 2025poster

This paper presents a groundbreaking approach - the first online automatic geometric calibration method for radar and camera systems. Given the significant data sparsity and measurement uncertainty in radar height data, achieving automatic calibration during system operation has long been a challeng…

2025

Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splatting

CVPR 2025poster

Bird's-eye view (BEV) perception has gained significant attention because it provides a unified representation to fuse multiple view images and enables a wide range of downstream autonomous driving tasks, such as forecasting and planning. Recent state-of-the-art models utilize projection-based metho…

Cited by 0SourcePDFScholar
2025

VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation

ICLR 2025poster

We study reward models for long-horizon manipulation by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Existing VIC methods face challenges in learning rewards for long-horizon tasks due to their lack of sub-stage awareness…

2025

What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

ICCV 2025poster

Understanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling…

2024

AED: Adaptable Error Detection for Few-shot Imitation Policy

NeurIPS 2024poster

We introduce a new task called Adaptable Error Detection (AED), which aims to identify behavior errors in few-shot imitation (FSI) policies based on visual observations in novel environments. The potential to cause serious damage to surrounding areas limits the application of FSI policies in real-wo…

2024

Action-slot: Visual Action-centric Representations for Multi-label Atomic Activity Recognition in Traffic Scenes

CVPR 2024poster

In this paper we study multi-label atomic activity recognition. Despite the notable progress in action recognition it is still challenging to recognize atomic activities due to a deficiency in holistic understanding of both multiple road users' motions and their contextual information. In this paper…

2024

RiskBench: A Scenario-based Benchmark for Risk Identification

ICRA 2024poster

Intelligent driving systems aim to achieve a zero-collision mobility experience, requiring interdisciplinary efforts to enhance safety performance. This work focuses on risk identification, the process of identifying and analyzing risks stemming from dynamic traffic participants and unexpected event…

Cited by 5SourcecodeScholar
2024

SKT-Hang: Hanging Everyday Objects via Object-Agnostic Semantic Keypoint Trajectory Generation

ICRA 2024poster

We study the problem of hanging a wide range of grasped objects on diverse supporting items. Hanging objects is a ubiquitous task that is encountered in numerous aspects of our everyday lives. However, both the objects and supporting items can exhibit substantial variations in their shapes and struc…

Cited by 0SourcecodeScholar
2023

Content Estimation Through Tactile Interactions with Deformable Containers

IROS 2023poster

Pouring snacks and moving containers with beverages are challenging for a service robot. To obtain accurate content properties for planning robotic motion, tactile sensing can provide information about the pressure distribution of the contact surface, which is not obvious by visual observation. In t…

Cited by 0SourceScholar
2023

Orbeez-SLAM: A Real-time Monocular Visual SLAM with ORB Features and NeRF-realized Mapping

ICRA 2023poster

A spatial AI that can perform complex tasks through visual signals and cooperate with humans is highly anticipated. To achieve this, we need a visual SLAM that easily adapts to new scenes without pre-training and generates dense maps for downstream tasks in real-time. None of the previous learning-b…

Cited by 137SourcecodeScholar
2023

SCONE: A Food Scooping Robot Learning Framework with Active Perception

CoRL 2023poster

Effectively scooping food items poses a substantial challenge for current robotic systems, due to the intricate states and diverse physical properties of food. To address this challenge, we believe in the importance of encoding food items into meaningful representations for effective food scooping.…

Cited by 12SourceScholar
2023

Shape-Aware Text-Driven Layered Video Editing

CVPR 2023poster

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV…

2022

Multimodal Object Detection via Probabilistic Ensembling

ECCV 2022poster

"Object detection with multimodal inputs can improve many safety-critical systems such as autonomous vehicles (AVs). Motivated by AVs that operate in both day and night, we study multimodal object detection with RGB and thermal cameras, since the latter provides much stronger object signatures under…

2022

Stage Conscious Attention Network (SCAN): A Demonstration-Conditioned Policy for Few-Shot Imitation

AAAI 2022technical

In few-shot imitation learning (FSIL), using behavioral cloning (BC) to solve unseen tasks with few expert demonstrations becomes a popular research direction. The following capabilities are essential in robotics applications: (1) Behaving in compound tasks that contain multiple stages. (2) Retrievi…

Cited by 4SourcePDFScholar
2021

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

ICML 2021oral

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive o…

Cited by 4467SourcePDFScholar
2020

Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks

ICRA 2020poster

To enable intelligent automated driving systems, a promising strategy is to understand how human drives and interacts with road users in complicated driving situations. In this paper, we propose a 3D-aware egocentric spatial-temporal interaction framework for automated driving applications. Graph co…

Cited by 78SourceScholar
2020

Who Make Drivers Stop? Towards Driver-centric Risk Assessment: Risk Object Identification via Causal Inference

IROS 2020poster

A significant amount of people die in road accidents due to driver errors. To reduce fatalities, developing intelligent driving systems assisting drivers to identify potential risks is in an urgent need. Risky situations are generally defined based on collision prediction in the existing works. Howe…

Cited by 63SourceScholar
2019

FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose Estimation From a Single Image

CVPR 2019poster

This paper proposes a method for head pose estimation from a single image. Previous methods often predict head poses through landmark or depth estimation and would require more computation than necessary. Our method is based on regression and feature aggregation. For having a compact model, we emplo…

Cited by 384PDFcodeScholar
2019

Grounding Human-To-Vehicle Advice for Self-Driving Vehicles

CVPR 2019poster

Recent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human actions, but they lack semantic understanding of image contents. This makes them brittle and potentially unsafe in situat…

Cited by 131PDFScholar
2019

Temporal Recurrent Networks for Online Action Detection

ICCV 2019poster

Most work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions…

Cited by 233PDFcodeScholar
2019

The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes

ICRA 2019poster

3D multi-object detection and tracking are crucial for traffic scene understanding. However, the community pays less attention to these areas due to the lack of a standardized benchmark dataset to advance the field. Moreover, existing datasets (e.g., KITTI [1]) do not provide sufficient data and lab…

Cited by 326SourceScholar
2018

Toward Driving Scene Understanding: A Dataset for Learning Driver Behavior and Causal Reasoning

CVPR 2018poster

Driving Scene understanding is a key ingredient for intelligent transportation systems. To achieve systems that can operate in a complex physical and social environment, they need to understand and learn how humans drive and interact with traffic scenes. We present the Honda Research Institute Drivi…

Cited by 406SourcePDFScholar