← Search

Qiangqiang Wu

13 accepted papers

2026

SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking

AAAI 2026technical

Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in

Cited by 0SourcePDFScholar
2026

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

AAAI 2026technical

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, t

Cited by 0SourcePDFScholar
2026

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

ICLR 2026poster

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to sele…

Cited by 0SourcecodeScholar
2025

DistinctAD: Distinctive Audio Description Generation in Contexts

CVPR 2025highlight

Audio Descriptions (ADs) aim to provide a narration of a movie in text form, describing non-dialogue-related narratives, such as characters, actions, or scene establishment. Automatic generation of ADs remains challenging due to: i) the domain gap between movie-AD data and existing data used to trai…

Cited by 2SourcePDFScholar
2025

MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence Learning

IJCAI 2025

Two-view correspondence learning is a key task in computer vision, which aims to establish reliable matching relationships for applications such as camera pose estimation and 3D reconstruction. However, existing methods have limitations in local geometric modeling and cross-stage information optimiz

Cited by 0SourcePDFScholar
2025

Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking

ICCV 2025poster

With the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial…

Cited by 0SourcePDFScholar
2024

Boosting 3D Single Object Tracking with 2D Matching Distillation and 3D Pre-training

ECCV 2024poster

"3D single object tracking (SOT) is an essential task in autonomous driving and robotics. However, learning robust 3D SOT trackers remains challenging due to the limited category-specific point cloud data and the inherent sparsity and incompleteness of LiDAR scans. To tackle these issues, we propose…

Cited by 3SourcePDFScholar
2024

Robust Zero-Shot Crowd Counting and Localization with Adaptive Resolution SAM

ECCV 2024poster

"The existing crowd counting models require extensive training data, which is time-consuming to annotate. To tackle this issue, we propose a simple yet effective crowd counting method by utilizing the Segment-Everything-Everywhere Model (SEEM), an adaptation of the Segmentation Anything Model (SAM),…

Cited by 3SourcePDFScholar
2023

DropMAE: Masked Autoencoders With Spatial-Attention Dropout for Tracking Tasks

CVPR 2023poster

In this paper, we study masked autoencoder (MAE) pretraining on videos for matching-based downstream tasks, including visual object tracking (VOT) and video object segmentation (VOS). A simple extension of MAE is to randomly mask out frame patches in videos and reconstruct the frame pixels. However,…

2023

TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer

ICCV 2023poster

In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caus…

Cited by 51PDFcodeScholar
2022

A New Framework for Multiple Deep Correlation Filters Based Object Tracking

ICASSP 2022accepted

In recent years, Correlation Filter (CF) based tracking methods using Convolutional Neural Network (CNN) features have achieved the state-of-the-art performance for object tracking. However, how to design an efficient deep CF based tracking method has not been well studied in the literature. To addr…

Cited by 0SourceScholar