← Search

Tsung-Han Wu

17 accepted papers

2026

SAM 3: Segment Anything with Concepts

ICLR 2026poster

We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both. Promptable Concept Segmentation (P…

Cited by 687SourcecodeScholar
2026

Search Arena: Analyzing Search-Augmented LLMs

ICLR 2026poster

Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in scale and narrow in scope, often constrained to static, single-turn, fact-check…

Cited by 0SourcecodeScholar
2025

Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling

NeurIPS 2025poster

Vision-Language Models (VLMs) excel at visual understanding but often suffer from visual hallucinations, where they generate descriptions of nonexistent objects, actions, or concepts, posing significant risks in safety-critical applications. Existing hallucination mitigation methods typically follow…

Cited by 0SourcecodeScholar
2025

LISAt: Language-Instructed Segmentation Assistant for Satellite Imagery

NeurIPS 2025poster

Segmentation models can recognize a pre-defined set of objects in images. However, segmentation models capable of "reasoning" over complex user queries that implicitly refer to multiple objects of interest remain underexplored, especially in the geospatial domain. Recent advances in "reasoning segme…

Cited by 0SourcecodeScholar
2025

Puzzled by Puzzles: When Vision-Language Models Can’t Take a Hint

EMNLP 2025

Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs). Unlike traditional image captioning or question answering tasks, rebus solving requires multimodal abstraction, symbol

2025

Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

ICLR 2025poster

Large Multimodal Models (LMMs) have made significant strides in visual question-answering for single images. Recent advancements like long-context LMMs have allowed them to ingest larger, or even multiple, images. However, the ability to process a large number of visual tokens does not guarantee eff…

2024

AED: Adaptable Error Detection for Few-shot Imitation Policy

NeurIPS 2024poster

We introduce a new task called Adaptable Error Detection (AED), which aims to identify behavior errors in few-shot imitation (FSI) policies based on visual observations in novel environments. The potential to cause serious damage to surrounding areas limits the application of FSI policies in real-wo…

2024

See Say and Segment: Teaching LMMs to Overcome False Premises

CVPR 2024poster

Current open-source Large Multimodal Models (LMMs) excel at tasks such as open-vocabulary language grounding and segmentation but can suffer under false premises when queries imply the existence of something that is not actually present in the image. We observe that existing methods that fine-tune a…

Cited by 20SourcePDFScholar
2024

Self-correcting LLM-controlled Diffusion Models

CVPR 2024poster

Text-to-image generation has witnessed significant progress with the advent of diffusion models. Despite the ability to generate photorealistic images current text-to-image diffusion models still often struggle to accurately interpret and follow complex input text prompts. In contrast to existing mo…

2024

WLST: Weak Labels Guided Self-training for Weakly-supervised Domain Adaptation on 3D Object Detection

ICRA 2024poster

In the field of domain adaptation (DA) on 3D object detection, most of the work is dedicated to unsupervised domain adaptation (UDA). Yet, without any target annotations, the performance gap between the UDA approaches and the fully-supervised approach is still noticeable, which is impractical for re…

Cited by 0SourcecodeScholar
2023

CrossDTR: Cross-view and Depth-guided Transformers for 3D Object Detection

ICRA 2023poster

To achieve accurate 3D object detection at a low cost for autonomous driving, many multi-camera methods have been proposed and solved the occlusion problem of monocular approaches. However, due to the lack of accurate estimated depth, existing multi-camera methods often generate multiple bounding bo…

Cited by 11SourcecodeScholar
2023

Revisiting Depth-guided Methods for Monocular 3D Object Detection by Hierarchical Balanced Depth

CoRL 2023poster

Monocular 3D object detection has seen significant advancements with the incorporation of depth information. However, there remains a considerable performance gap compared to LiDAR-based methods, largely due to inaccurate depth estimation. We argue that this issue stems from the commonly used pixel-…

Cited by 1SourceScholar
2022

D2ADA: Dynamic Density-Aware Active Domain Adaptation for Semantic Segmentation

ECCV 2022poster

"In the field of domain adaptation, a trade-off exists between the model performance and the number of target domain annotations. Active learning, maximizing model performance with few informative labeled data, comes in handy for such a scenario. In this work, we present D2ADA, a general active doma…

2022

MonoDTR: Monocular 3D Object Detection With Depth-Aware Transformer

CVPR 2022poster

Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the additional computational burden and achieve limited performance caused by inacc…

Cited by 214PDFcodeScholar
2021

Again-VC: A One-Shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization

ICASSP 2021accepted

Recently, voice conversion (VC) has been widely studied. Many VC systems use disentangle-based learning techniques to separate the speaker and the linguistic content information from a speech signal. Subsequently, they convert the voice by changing the speaker information to that of the target speak…

Cited by 0SourceScholar
2021

ReDAL: Region-Based and Diversity-Aware Active Learning for Point Cloud Semantic Segmentation

ICCV 2021poster

Despite the success of deep learning on supervised point cloud semantic segmentation, obtaining large-scale point-by-point manual annotations is still a significant challenge. To reduce the huge annotation burden, we propose a Region-based and Diversity-aware Active Learning (ReDAL), a general frame…

Cited by 97PDFcodeScholar
2021

S3: Learnable Sparse Signal Superdensity for Guided Depth Estimation

CVPR 2021poster

Dense depth estimation plays a key role in multiple applications such as robotics, 3D reconstruction, and augmented reality. While sparse signal, e.g., LiDAR and Radar, has been leveraged as guidance for enhancing dense depth estimation, the improvement is limited due to its low density and imbalanc…

Cited by 22PDFScholar