← Search

Winston H. Hsu

38 accepted papers

2026

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

AAAI 2026technical

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances, resulting in frequent manipulation failures. We propose Afforda

Cited by 0SourcePDFScholar
2025

Attention Tracker: Detecting Prompt Injection Attacks in LLMs

NAACL 2025findings

Large Language Models (LLMs) have revolutionized various domains but remain vulnerable to prompt injection attacks, where malicious inputs manipulate the model into ignoring original instructions and executing designated action. In this paper, we investigate the underlying mechanisms of these attack…

2025

HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics

ICCV 2025poster

Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propo…

2025

Improving Generalization Ability for 3D Object Detection by Learning Sparsity-Invariant Features

ICRA 2025

In autonomous driving, 3D object detection is essential for accurately identifying and tracking objects. Despite the continuous development of various technologies for this task, a significant drawback is observed in most of them—they experience substantial performance degradation when detecting obj

Cited by 1SourcecodeScholar
2025

MovieCORE: COgnitive REasoning in Movies

EMNLP 2025

This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes questions that engage System-2 thinking while remaining specific

2025

VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation

ICLR 2025poster

We study reward models for long-horizon manipulation by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Existing VIC methods face challenges in learning rewards for long-horizon tasks due to their lack of sub-stage awareness…

2024

AED: Adaptable Error Detection for Few-shot Imitation Policy

NeurIPS 2024poster

We introduce a new task called Adaptable Error Detection (AED), which aims to identify behavior errors in few-shot imitation (FSI) policies based on visual observations in novel environments. The potential to cause serious damage to surrounding areas limits the application of FSI policies in real-wo…

2024

Context-Aware Replanning with Pre-Explored Semantic Map for Object Navigation

CoRL 2024poster

Pre-explored Semantic Map, constructed through prior exploration using visual language models (VLMs), has proven effective as a foundational element for training-free robotic applications. However, existing approaches assume the map's accuracy and do not provide effective mechanisms for revising dec…

Cited by 0SourceScholar
2024

Distribution Discrepancy and Feature Heterogeneity for Active 3D Object Detection

CoRL 2024poster

LiDAR-based 3D object detection is a critical technology for the development of autonomous driving and robotics. However, the high cost of data annotation limits its advancement. We propose a novel and effective active learning (AL) method called Distribution Discrepancy and Feature Heterogeneity (D…

Cited by 0SourcecodeScholar
2024

Enhancing Sustainable Urban Mobility Prediction with Telecom Data: A Spatio-Temporal Framework Approach

IJCAI 2024poster

Traditional traffic prediction, limited by the scope of sensor data, falls short in comprehensive traffic management. Mobile networks offer a promising alternative using network activity counts, but these lack crucial directionality. Thus, we present the TeltoMob dataset, featuring undirected teleco…

2024

TelTrans: Applying Multi-Type Telecom Data to Transportation Evaluation and Prediction via Multifaceted Graph Modeling

AAAI 2024technical

To address the limitations of traffic prediction from location-bound detectors, we present Geographical Cellular Traffic (GCT) flow, a novel data source that leverages the extensive coverage of cellular traffic to capture mobility patterns. Our extensive analysis validates its potential for transpor…

2024

Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses

EMNLP 2024finding

Large language models (LLMs) equipped with chain-of-thoughts (CoT) prompting have shown significant multi-step reasoning capabilities in factual content like mathematics, commonsense, and logic. However, their performance in narrative reasoning, which demands greater abstraction capabilities, remain…

2024

WLST: Weak Labels Guided Self-training for Weakly-supervised Domain Adaptation on 3D Object Detection

ICRA 2024poster

In the field of domain adaptation (DA) on 3D object detection, most of the work is dedicated to unsupervised domain adaptation (UDA). Yet, without any target annotations, the performance gap between the UDA approaches and the fully-supervised approach is still noticeable, which is impractical for re…

Cited by 0SourcecodeScholar
2023

BIRD-PCC: Bi-Directional Range Image-Based Deep Lidar Point Cloud Compression

ICASSP 2023accepted

The large amount of data collected by LiDAR sensors brings the issue of LiDAR point cloud compression (PCC). Previous works on LiDAR PCC have used range image representations and followed the predictive coding paradigm to create a basic prototype of a coding framework. However, their prediction meth…

Cited by 0SourceScholar
2023

CFVS: Coarse-to-Fine Visual Servoing for 6-DoF Object-Agnostic Peg-In-Hole Assembly

ICRA 2023poster

Robotic peg-in-hole assembly remains a challenging task due to its high accuracy demand. Previous work tends to simplify the problem by restricting the degree of freedom of the end-effector, or limiting the distance between the target and the initial pose position, which prevents them from being dep…

Cited by 16SourceScholar
2023

Coarse-to-Fine Point Cloud Registration with SE(3)-Equivariant Representations

ICRA 2023poster

Point cloud registration is a crucial problem in computer vision and robotics. Existing methods either rely on matching local geometric features, which are sensitive to the pose differences, or leverage global shapes, which leads to inconsistency when facing distribution variances such as partial ov…

Cited by 18SourcecodeScholar
2023

CrossDTR: Cross-view and Depth-guided Transformers for 3D Object Detection

ICRA 2023poster

To achieve accurate 3D object detection at a low cost for autonomous driving, many multi-camera methods have been proposed and solved the occlusion problem of monocular approaches. However, due to the lack of accurate estimated depth, existing multi-camera methods often generate multiple bounding bo…

Cited by 11SourcecodeScholar
2023

Orbeez-SLAM: A Real-time Monocular Visual SLAM with ORB Features and NeRF-realized Mapping

ICRA 2023poster

A spatial AI that can perform complex tasks through visual signals and cooperate with humans is highly anticipated. To achieve this, we need a visual SLAM that easily adapts to new scenes without pre-training and generates dense maps for downstream tasks in real-time. None of the previous learning-b…

Cited by 137SourcecodeScholar
2023

Revisiting Depth-guided Methods for Monocular 3D Object Detection by Hierarchical Balanced Depth

CoRL 2023poster

Monocular 3D object detection has seen significant advancements with the incorporation of depth information. However, there remains a considerable performance gap compared to LiDAR-based methods, largely due to inaccurate depth estimation. We argue that this issue stems from the commonly used pixel-…

Cited by 1SourceScholar
2022

D2ADA: Dynamic Density-Aware Active Domain Adaptation for Semantic Segmentation

ECCV 2022poster

"In the field of domain adaptation, a trade-off exists between the model performance and the number of target domain annotations. Active learning, maximizing model performance with few informative labeled data, comes in handy for such a scenario. In this work, we present D2ADA, a general active doma…

2022

MonoDTR: Monocular 3D Object Detection With Depth-Aware Transformer

CVPR 2022poster

Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the additional computational burden and achieve limited performance caused by inacc…

Cited by 214PDFcodeScholar
2022

Stage Conscious Attention Network (SCAN): A Demonstration-Conditioned Policy for Few-Shot Imitation

AAAI 2022technical

In few-shot imitation learning (FSIL), using behavioral cloning (BC) to solve unseen tasks with few expert demonstrations becomes a popular research direction. The following capabilities are essential in robotics applications: (1) Behaving in compound tasks that contain multiple stages. (2) Retrievi…

Cited by 4SourcePDFScholar
2021

ODIP: Towards Automatic Adaptation for Object Detection by Interactive Perception

IROS 2021poster

Object detection plays a deep role in visual systems by identifying instances for downstream algorithms. In industrial scenarios, however, a slight change in manufacturing systems would lead to costly data re-collection and human annotation processes to re-train models. Existing solutions such as se…

Cited by 0SourceScholar
2021

ReDAL: Region-Based and Diversity-Aware Active Learning for Point Cloud Semantic Segmentation

ICCV 2021poster

Despite the success of deep learning on supervised point cloud semantic segmentation, obtaining large-scale point-by-point manual annotations is still a significant challenge. To reduce the huge annotation burden, we propose a Region-based and Diversity-aware Active Learning (ReDAL), a general frame…

Cited by 97PDFcodeScholar
2021

Role Aware Multi-Party Dialogue Question Answering

ICASSP 2021accepted

Multi-party dialogue question answering (MPDQA) is an emerging topic in speech and language processing where the goal is to answer the questions according to the multiparty conversations. Different from conventional QA, which assumes a single speaker (writer) and general listeners (readers), MPDQA i…

Cited by 0SourceScholar
2021

S3: Learnable Sparse Signal Superdensity for Guided Depth Estimation

CVPR 2021poster

Dense depth estimation plays a key role in multiple applications such as robotics, 3D reconstruction, and augmented reality. While sparse signal, e.g., LiDAR and Radar, has been leveraged as guidance for enhancing dense depth estimation, the improvement is limited due to its low density and imbalanc…

Cited by 22PDFScholar
2020

Video Question Generation via Semantic Rich Cross-Modal Self-Attention Networks Learning

ICASSP 2020accepted

We introduce a novel task, Video Question Generation (Video QG). A Video QG model automatically generates questions given a video clip and its corresponding dialogues. Video QG requires a range of skills - sentence comprehension, temporal relation, the interplay between vision and language, and the…

Cited by 0SourceScholar
2019

Audio Feature Generation for Missing Modality Problem in Video Action Recognition

ICASSP 2019accepted

Despite the recent success of multi-modal action recognition in videos, in reality, we usually confront the situation that some data are not available beforehand, especially for multi-modal data. For example, while vision and audio data are required to address the multi-modal action recognition, aud…

Cited by 0SourceScholar
2018

Super-Identity Convolutional Neural Network for Face Hallucination

ECCV 2018poster

Face hallucination is a generative task to super-resolve the facial image with low resolution while human perception of face heavily relies on identity information. However, previous face hallucination approaches largely ignore facial identity recovery. This paper proposes Super-Identity Convolution…

Cited by 162SourcePDFScholar
2017

Drone-Based Object Counting by Spatially Regularized Regional Proposal Network

ICCV 2017poster

Existing counting methods often adopt regression-based approaches and cannot precisely localize the target objects, which hinders the further analysis (e.g., high-level understanding and fine-grained classification). In addition, most of prior work mainly focus on counting objects in static environm…

Cited by 535PDFScholar
2015

Enhancing sparse voice annotation for semantic retrieval of personal photos by continuous space word representations

ICASSP 2015accepted

It is very attractive for the user to retrieve photos from a huge collection using high-level personal queries (e.q. uncle Bill's house), but technically very challenging. The previous work proposed a set of approaches to achieve the goal assuming only 30% of the photos are annotated by sparse spoke…

Cited by 0SourceScholar
2015

Identify Visual Human Signature in community via wearable camera

ICASSP 2015accepted

With the increasing popularity of wearable devices, information becomes much easily available. However, personal information sharing still poses great challenges because of privacy issues. We propose an idea of Visual Human Signature (VHS) which can represent each person uniquely even captured in di…

Cited by 0SourceScholar
2015

Scalable Object Detection by Filter Compression With Regularized Sparse Coding

CVPR 2015poster

For practical applications, an object detection system requires huge number of classes to meet real world needs. Many successful object detection systems use part-based model which trains several filters (classifiers) for each class to perform multiclass object detection. However, these methods have…

Cited by 8SourcePDFScholar