← Search

Hsiang-Wei Huang

7 accepted papers

2026

CLEP: Contrastive Language-Pose Pretraining

CVPR 2026

Aligning natural language descriptions with precise 3D human poses remains a big challenge due to the scarcity of effective pose representation mechanisms and large-scale, semantically rich datasets. To overcome these limitations, we first introduce **CLEP-2M**, the largest 3D pose-language dataset

Cited by 0SourceScholar
2025

Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

ICCV 2025poster

Unlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research,…

Cited by 0SourcePDFScholar
2025

MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking

ICASSP 2025accepted

In the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusion…

Cited by 0SourceScholar
2025

ToSA: Token Merging with Spatial Awareness

IROS 2025

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual token’s feature similarity for token merging, overlooking the potential of integrating spatial information, which can ser

Cited by 5SourcecodeScholar
2025

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

CVPR 2025poster

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and…

Cited by 0SourcePDFScholar
2024

A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos

ICASSP 2024accepted

Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counti…

Cited by 0SourceScholar
2024

RT-Pose: A 4D Radar-Tensor based 3D Human Pose Estimation and Localization Benchmark

ECCV 2024poster

"Traditional methods for human localization and pose estimation (HPE), which mainly rely on RGB images as an input modality, confront substantial limitations in real-world applications due to privacy concerns. In contrast, radar-based HPE methods emerge as a promising alternative, characterized by d…

Cited by 5SourcePDFScholar