← Search

Fu-Jen Chu

13 accepted papers

2026

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

CVPR 2026

Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physical-world applications that require multi-frame reasoning. In this paper, we propose a framework to equip MLLMs with mul

Cited by 0SourcecodeScholar
2025

HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models

CVPR 2025poster

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partia…

Cited by 1SourcePDFScholar
2025

OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB

IROS 2025

To address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive compa

Cited by 1SourceScholar
2024

"Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos"

ECCV 2024oral

"Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tasks. The problem presents significant challenges due to the need for comprehensi…

Cited by 2SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2023

Relational Space-Time Query in Long-Form Videos

CVPR 2023highlight

Egocentric videos are often available in the form of uninterrupted, uncurated long videos capturing the camera wearers' daily life activities.Understanding these videos requires models to be able to reason about activities, objects, and their interactions. However, current video benchmarks study the…

Cited by 14SourcePDFScholar
2020

Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object Grasping

ICRA 2020poster

A segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is trained on synthetic data with 6 classes of primitive shapes generated by a simulation engine. Each primitive shape is d…

Cited by 54SourceScholar
2019

Learning Affordance Segmentation for Real-World Robotic Manipulation via Synthetic Images

RA-L 2019

This letter presents a deep learning framework to predict the affordances of object parts for robotic manipulation. The framework segments affordance maps by jointly detecting and localizing candidate regions within an image. Rather than requiring annotated real-world images, the framework learns fr

Cited by 61SourceScholar
2019

Toward Affordance Detection and Ranking on Novel Objects for Real-World Robotic Manipulation

RA-L 2019

This letter presents a framework to detect and rank affordances of novel objects to assist with robotic manipulation tasks. The framework segments the affordance map of unseen objects using region-based affordance segmentation. Detected affordances define an initial state from which to generate acti

Cited by 40SourceScholar
2018

Hands-Free Assistive Manipulator Using Augmented Reality and Tongue Drive System

IROS 2018poster

A human-in-the-loop system is proposed to enable hands-free collaborative manipulation for people with physical disabilities. Studies show that the cognitive burden of interfacing with a robotic assistant decreases with increased robot autonomy. Incorporating modern advances in perception with augme…

Cited by 8SourceScholar