← Search

Guijian Tang

4 accepted papers

2026

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

CVPR 2026

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals into a single attention space. We observe diluted attention a

Cited by 0SourceScholar
2026

RAG-TP: A General Framework for Vehicle Trajectory Prediction via Retrieval-Augmented Generation

CVPR 2026

Vehicle trajectory prediction is critical for safe and efficient autonomous driving. However, its generalization and scalability are hindered by heavy reliance on real-time, online priors. To break this bottleneck, we introduce RAG-TP, a framework reframing the problem from relying on uncertain onli

Cited by 0SourceScholar
2026

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

CVPR 2026

Video reasoning has advanced with large multimodal models (LMMs), yet their inference is often a single pass that returns an answer without verifying whether the reasoning is evidence-aligned. We introduce **Reinforce to Learn, Elect to Reason (RLER)**, a dual paradigm that decouples learning to pro

Cited by 0SourceScholar
2026

Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucination

CVPR 2026

Hallucination limits the reliability of multimodal large language models (MLLMs), and it is particularly damaging in video where errors manifest as distorted narratives rather than single-frame mistakes. We introduce a frame-first study of **Chimera Hallucination**: model stitches visual segments th

Cited by 0SourceScholar