← Search

Hung-Ting Su

17 accepted papers

2026

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

AAAI 2026technical

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances, resulting in frequent manipulation failures. We propose Afforda

Cited by 0SourcePDFScholar
2025

HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics

ICCV 2025poster

Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propo…

2025

MovieCORE: COgnitive REasoning in Movies

EMNLP 2025

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific

2024

AED: Adaptable Error Detection for Few-shot Imitation Policy

NeurIPS 2024poster

We introduce a new task called Adaptable Error Detection (AED), which aims to identify behavior errors in few-shot imitation (FSI) policies based on visual observations in novel environments. The potential to cause serious damage to surrounding areas limits the application of FSI policies in real-wo…

2024

Context-Aware Replanning with Pre-Explored Semantic Map for Object Navigation

CoRL 2024poster

Pre-explored Semantic Map, constructed through prior exploration using visual language models (VLMs), has proven effective as a foundational element for training-free robotic applications. However, existing approaches assume the map's accuracy and do not provide effective mechanisms for revising dec…

Cited by 0SourceScholar
2024

Enhancing Sustainable Urban Mobility Prediction with Telecom Data: A Spatio-Temporal Framework Approach

IJCAI 2024poster

Traditional traffic prediction, limited by the scope of sensor data, falls short in comprehensive traffic management. Mobile networks offer a promising alternative using network activity counts, but these lack crucial directionality. Thus, we present the TeltoMob dataset, featuring undirected teleco…

2024

TelTrans: Applying Multi-Type Telecom Data to Transportation Evaluation and Prediction via Multifaceted Graph Modeling

AAAI 2024technical

To address the limitations of traffic prediction from location-bound detectors, we present Geographical Cellular Traffic (GCT) flow, a novel data source that leverages the extensive coverage of cellular traffic to capture mobility patterns. Our extensive analysis validates its potential for transpor…

2024

Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses

EMNLP 2024finding

Large language models (LLMs) equipped with chain-of-thoughts (CoT) prompting have shown significant multi-step reasoning capabilities in factual content like mathematics, commonsense, and logic. However, their performance in narrative reasoning, which demands greater abstraction capabilities, remain…

2023

BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression

ICASSP 2023accepted

The large amount of data collected by LiDAR sensors brings the issue of LiDAR point cloud compression (PCC). Previous works on LiDAR PCC have used range image representations and followed the predictive coding paradigm to create a basic prototype of a coding framework. However, their prediction meth…

Cited by 0SourceScholar
2022

MonoDTR: Monocular 3D Object Detection With Depth-Aware Transformer

CVPR 2022poster

Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the additional computational burden and achieve limited performance caused by inacc…

Cited by 214PDFcodeScholar
2022

Stage Conscious Attention Network (SCAN): A Demonstration-Conditioned Policy for Few-Shot Imitation

AAAI 2022technical

In few-shot imitation learning (FSIL), using behavioral cloning (BC) to solve unseen tasks with few expert demonstrations becomes a popular research direction. The following capabilities are essential in robotics applications: (1) Behaving in compound tasks that contain multiple stages. (2) Retrievi…

Cited by 4SourcePDFScholar
2021

OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding

NAACL 2021long

To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded objects. However, current VG works are limited in working environments, such as offices and warehouses, where objects ar…

2021

ReDAL: Region-Based and Diversity-Aware Active Learning for Point Cloud Semantic Segmentation

ICCV 2021poster

Despite the success of deep learning on supervised point cloud semantic segmentation, obtaining large-scale point-by-point manual annotations is still a significant challenge. To reduce the huge annotation burden, we propose a Region-based and Diversity-aware Active Learning (ReDAL), a general frame…

Cited by 97PDFcodeScholar
2021

Role Aware Multi-Party Dialogue Question Answering

ICASSP 2021accepted

Multi-party dialogue question answering (MPDQA) is an emerging topic in speech and language processing where the goal is to answer the questions according to the multiparty conversations. Different from conventional QA, which assumes a single speaker (writer) and general listeners (readers), MPDQA i…

Cited by 0SourceScholar
2021

S3: Learnable Sparse Signal Superdensity for Guided Depth Estimation

CVPR 2021poster

Dense depth estimation plays a key role in multiple applications such as robotics, 3D reconstruction, and augmented reality. While sparse signal, e.g., LiDAR and Radar, has been leveraged as guidance for enhancing dense depth estimation, the improvement is limited due to its low density and imbalanc…

Cited by 22PDFScholar
2020

GDN: A Coarse-To-Fine (C2F) Representation for End-To-End 6-DoF Grasp Detection

CoRL 2020

We proposed an end-to-end grasp detection network, Grasp Detection Network (GDN), cooperated with a novel coarse-to-fine (C2F) grasp representation design to detect diverse and accurate 6-DoF grasps based on point clouds. Compared to previous two-stage approaches which sample and evaluate multiple g

Cited by 0SourcePDFScholar
2020

Video Question Generation via Semantic Rich Cross-Modal Self-Attention Networks Learning

ICASSP 2020accepted

We introduce a novel task, Video Question Generation (Video QG). A Video QG model automatically generates questions given a video clip and its corresponding dialogues. Video QG requires a range of skills - sentence comprehension, temporal relation, the interplay between vision and language, and the…

Cited by 0SourceScholar