← Search

Jia-Fong Yeh

11 accepted papers

2026

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

AAAI 2026technical

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances, resulting in frequent manipulation failures. We propose Afforda

Cited by 0SourcePDFScholar
2025

HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics

ICCV 2025poster

Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propo…

2025

MovieCORE: COgnitive REasoning in Movies

EMNLP 2025

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific

2025

VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation

ICLR 2025poster

We study reward models for long-horizon manipulation by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Existing VIC methods face challenges in learning rewards for long-horizon tasks due to their lack of sub-stage awareness…

2024

AED: Adaptable Error Detection for Few-shot Imitation Policy

NeurIPS 2024poster

We introduce a new task called Adaptable Error Detection (AED), which aims to identify behavior errors in few-shot imitation (FSI) policies based on visual observations in novel environments. The potential to cause serious damage to surrounding areas limits the application of FSI policies in real-wo…

2024

Context-Aware Replanning with Pre-Explored Semantic Map for Object Navigation

CoRL 2024poster

Pre-explored Semantic Map, constructed through prior exploration using visual language models (VLMs), has proven effective as a foundational element for training-free robotic applications. However, existing approaches assume the map's accuracy and do not provide effective mechanisms for revising dec…

Cited by 0SourceScholar
2024

Distribution Discrepancy and Feature Heterogeneity for Active 3D Object Detection

CoRL 2024poster

LiDAR-based 3D object detection is a critical technology for the development of autonomous driving and robotics. However, the high cost of data annotation limits its advancement. We propose a novel and effective active learning (AL) method called Distribution Discrepancy and Feature Heterogeneity (D…

Cited by 0SourcecodeScholar
2023

BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression

ICASSP 2023accepted

The large amount of data collected by LiDAR sensors brings the issue of LiDAR point cloud compression (PCC). Previous works on LiDAR PCC have used range image representations and followed the predictive coding paradigm to create a basic prototype of a coding framework. However, their prediction meth…

Cited by 0SourceScholar
2023

Orbeez-SLAM: A Real-time Monocular Visual SLAM with ORB Features and NeRF-realized Mapping

ICRA 2023poster

A spatial AI that can perform complex tasks through visual signals and cooperate with humans is highly anticipated. To achieve this, we need a visual SLAM that easily adapts to new scenes without pre-training and generates dense maps for downstream tasks in real-time. None of the previous learning-b…

Cited by 137SourcecodeScholar
2022

Stage Conscious Attention Network (SCAN): A Demonstration-Conditioned Policy for Few-Shot Imitation

AAAI 2022technical

In few-shot imitation learning (FSIL), using behavioral cloning (BC) to solve unseen tasks with few expert demonstrations becomes a popular research direction. The following capabilities are essential in robotics applications: (1) Behaving in compound tasks that contain multiple stages. (2) Retrievi…

Cited by 4SourcePDFScholar
2021

Role Aware Multi-Party Dialogue Question Answering

ICASSP 2021accepted

Multi-party dialogue question answering (MPDQA) is an emerging topic in speech and language processing where the goal is to answer the questions according to the multiparty conversations. Different from conventional QA, which assumes a single speaker (writer) and general listeners (readers), MPDQA i…

Cited by 0SourceScholar