← Search

Huibin Tan

11 accepted papers

2026

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

CVPR 2026

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals into a single attention space. We observe diluted attention a

Cited by 0SourceScholar
2026

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

CVPR 2026

Video reasoning has advanced with large multimodal models (LMMs), yet their inference is often a single pass that returns an answer without verifying whether the reasoning is evidence-aligned. We introduce **Reinforce to Learn, Elect to Reason (RLER)**, a dual paradigm that decouples learning to pro

Cited by 0SourceScholar
2026

Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination

CVPR 2026

Hallucination limits the reliability of multimodal large language models (MLLMs), and it is particularly damaging in video where errors manifest as distorted narratives rather than single-frame mistakes. We introduce a frame-first study of **Chimera Hallucination**: model stitches visual segments th

Cited by 0SourceScholar
2025

UniIVFT: Towards a Unified Framework for Infrared-Visible Fusion and Translation

ICASSP 2025accepted

Infrared-visible image fusion (IVF) and infrared-to-visible image translation (I2V) are two closely related tasks in multimodal image processing, both aimed at combining or transforming infrared and visible modalities to enhance image information content. Existing methods typically focus on either f…

Cited by 0SourceScholar
2025

Wave-wise Discriminative Tracking by Phase-Amplitude Separation, Augmentation and Mixture

IJCAI 2025

Distinguishing key features in complex visual tasks is challenging. A novel approach treats image patches (tokens) as waves. By using both phase and amplitude, it captures richer semantics and specific invariances compared to pixel-based methods, and allows for feature fusion across regions for a ho

Cited by 0SourcePDFScholar
2024

Radar Recognition in the Wild: Enhancing Radar Emitter Recognition through Auto-Correlation Model-Agnostic Meta Learning

ICASSP 2024accepted

In Electronic Support Measure (ESM) systems, the recognition of radar emitters stands as a pivotal yet intricate task. The complex electromagnetic environments, however, often hinders the collection of clean radar signal data, and results in data with different noise levels. Consequently, formulatin…

Cited by 0SourceScholar
2024

Sequential Fusion Based Multi-Granularity Consistency for Space-Time Transformer Tracking

AAAI 2024technical

Regarded as a template-matching task for a long time, visual object tracking has witnessed significant progress in space-wise exploration. However, since tracking is performed on videos with substantial time-wise information, it is important to simultaneously mine the temporal contexts which have no…

Cited by 7SourcePDFScholar
2023

Decomposition, Interaction, Reconstruction Meets Global Context Learning In Visual Tracking

ICASSP 2023accepted

Tensor decomposition and reconstruction attention is a promising global context learning approach because it can remain efficient while avoiding feature compression. To exploit its potential even further in visual tracking, we redesign a 3D tensor modeling paradigm, namely tensor Decomposition, Inte…

Cited by 0SourceScholar
2023

Enhanced Dcf Tracker Regularized by Reliable Sample Construction

ICASSP 2023accepted

Discriminative correlation filter (DCF) is a highly efficient tracking technique using the circulant shifted samples of search images to update the template, so the reliability of input samples determines template quality. In this paper, we rethink the reliability problem of input samples in advance…

Cited by 0SourceScholar
2023

Progressive Perception Learning for Distribution Modulation in Siamese Tracking

ICASSP 2023accepted

We explore an innovative view on distribution modulation to boost Siamese trackers. Specially, we observed two cases of possible distribution inconsistency in Siamese tracking: 1) Two branches with different sizes may be in different distribution ranges after a shared backbone (including BN layers).…

Cited by 0SourceScholar