← Search

Cheng-Yen Yang

6 accepted papers

2025

MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking

ICASSP 2025accepted

In the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusion…

Cited by 0SourceScholar
2025

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

NeurIPS 2025poster

Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. Nevertheless, the coarse-to-fine methodology inherent in VAR results in exponenti…

Cited by 0SourcecodeScholar
2025

ToSA: Token Merging with Spatial Awareness

IROS 2025

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual token’s feature similarity for token merging, overlooking the potential of integrating spatial information, which can ser

Cited by 5SourcecodeScholar
2025

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

CVPR 2025poster

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and…

Cited by 0SourcePDFScholar
2024

2D Human Pose Estimation Calibration and Keypoint Visibility Classification

ICASSP 2024accepted

The confidence scores of 2D pose estimation are widely utilized in various fields, including multi-view 3D human pose estimation, skeleton-based human tracking, human action recognition, human re-identification, etc. Despite widespread use, confidence scores from 2D pose estimation methods are unrel…

Cited by 0SourceScholar
2024

A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos

ICASSP 2024accepted

Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counti…

Cited by 0SourceScholar