← Search

Teng Wang

35 accepted papers

2026

AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos

CVPR 2026

Long-term language-guided referring in fixed-view videos is challenging: the referent may be occluded or leave the scene for long intervals and later re-enter, while framewise referring pipelines drift as re-identification (ReID) becomes unreliable. AR2-4FV leverages background stability for long-te

Cited by 0SourceScholar
2026

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

CVPR 2026

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To fill this gap, we propose AudioStory, a unified framework that integrates large language models (LL

Cited by 0SourcecodeScholar
2026

CP-Router: An Uncertainty-Aware Router Between LLM and LRM

AAAI 2026technical

Recent advances in large reasoning models (LRMs) have significantly enhanced long-chain reasoning capabilities over standard large language models (LLMs). However, LRMs often produce unnecessarily lengthy outputs even for simple queries, leading to inefficiencies or even accuracy degradation compare

Cited by 0SourcePDFScholar
2026

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

ICML 2026poster

Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progres…

Cited by 0SourceScholar
2026

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

ICML 2026poster

Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This paradigm overlooks the rich dependencies inherent in realistic visual streams, where information is distributed across temporal sequences rather than c…

Cited by 0SourceScholar
2026

FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning

ICLR 2026poster

Group relative policy optimization (GRPO) has demonstrated significant potential in improving the reasoning capabilities of large language models (LLMs) via reinforcement learning. However, its practical deployment is impeded by an excessively slow training process, primarily attributed to the compu…

Cited by 0SourcecodeScholar
2026

First Learn, Then Review: Human-Like Continual Learning for Cross-View Geo-Localization with Limited Field of View

AAAI 2026technical

This paper addresses cross-view geo-localization in real-world scenarios, where the field-of-view (FoV) is restricted and the orientation is unknown for ground-view images. This task is extremely challenging due to the huge domain gap. Existing methods typically treat tasks with different FoVs as in

Cited by 0SourcePDFScholar
2026

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

AAAI 2026technical

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the complex and diverse nature of real-world audio-visual events in videos. To bridge this

Cited by 0SourcePDFScholar
2026

RaCo-SLAM: A Physics-Informed 4D Radar SLAM with Co-Visibility Consistency Factor

ICRA 2026poster

Robust all-weather localization is a critical capability for autonomous systems. While 4D mmWave radar offers superior resilience to adverse environmental conditions compared to LiDAR and cameras, its application in high-precision Simultaneous Localization and Mapping (SLAM) is hindered by significa…

Cited by 0codeScholar
2026

TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

CVPR 2026

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal large language models (MLLMs) excel at various video understanding tasks, the rec

Cited by 0SourceScholar
2026

UltraHiT: A Hierarchical Transformer Architecture for Generalizable Internal Carotid Artery Robotic Ultrasonography

ICRA 2026poster

Carotid ultrasound is crucial for the assessment of cerebrovascular health, particularly the internal carotid artery (ICA). While previous research has explored automating carotid ultrasound, none has tackled the challenging ICA. This is primarily due to its deep location, tortuous course, and signi…

2025

BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving

ACL 2025long

LLMs exhibit advanced reasoning capabilities, offering the potential to transform natural language questions into mathematical models. However, existing open-source datasets in operations research domain lack detailed annotations of the modeling process, such as variable definitions, focusing solely…

2025

Boosting Efficient Reinforcement Learning for Vision-and-Language Navigation With Open-Sourced LLM

RA-L 2025

Vision-and-Language Navigation (VLN) requires an agent to navigate in photo-realistic environments based on language instructions. Existing methods typically employ imitation learning to train agents. However, approaches based on recurrent neural networks suffer from poor generalization, while trans

Cited by 13SourceScholar
2025

Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal Models

IJCAI 2025

Dynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding bu

Cited by 0SourcePDFScholar
2025

EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow Estimation

CVPR 2025poster

Recent learning-based methods for event-based optical flow estimation utilize cost volumes for pixel matching but suffer from redundant computations and limited scalability to higher resolutions for flow refinement. In this work, we take advantage of the complementarity between temporally dense feat…

Cited by 0SourcePDFScholar
2025

GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers

ICCV 2025poster

The synergy between generative and discriminative models receives growing attention. While discriminative Contrastive Language-Image Pre-Training (CLIP) excels in high-level semantics, it struggles with perceiving fine-grained visual details. Generally, to enhance representations, generative models…

Cited by 0SourcePDFScholar
2025

Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency

IJCAI 2025

The rapid advancement of large language models (LLMs) has led to the widespread adoption of video-language models (VLMs) across various domains. However, VLMs are often hindered by their limited semantic discrimination capability, exacerbated by the limited diversity and biased sample distribution o

Cited by 0SourcePDFScholar
2025

ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

ICLR 2025poster

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLM…

Cited by 2SourcePDFScholar
2025

Large Language Models are good multi-lingual learners : When LLMs meet cross-lingual prompts

COLING 2025main

With the advent of Large Language Models (LLMs), generating rule-based data for real-world applications has become more accessible. Due to the inherent ambiguity of natural language and the complexity of rule sets, especially in long contexts, LLMs often struggle to follow all specified rules, frequ…

2025

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

CVPR 2025poster

Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal vide…

2025

Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models

ICLR 2025spotlight

Multimodal Large Language Models (MLLMs) exhibit promising advancements across various tasks, yet they still encounter significant trustworthiness issues. Prior studies apply Split Conformal Prediction (SCP) in language modeling to construct prediction sets with statistical guarantees. However, thes…

Cited by 6SourcePDFScholar
2025

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

EMNLP 2025

Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compression strategies apply a fixed compression ratio, ignoring the variability in sem

2024

Efficient Joint Rectification of Photometric and Geometric Distortions in Document Images

ICASSP 2024accepted

Document images captured with cameras often exhibit photometric and geometric distortions. Here, we propose a novel learning-based approach for efficient joint rectification of document images. Inspired by the strong correlation between visual shadows and physical deformations, we design a shared en…

Cited by 0SourceScholar
2024

Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models

ECCV 2024poster

"Large vision-language models (LVLMs) have shown promising performance on a variety of vision-language tasks. However, they remain susceptible to hallucinations, generating outputs misaligned with visual content or instructions. While various mitigation strategies have been proposed, they often negl…

Cited by 7SourcePDFScholar
2023

$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation

ICML 2023poster

Foundation models have achieved great advances in multi-task learning with a unified interface of unimodal and multimodal tasks. However, the potential of such multi-task learners has not been exploited during transfer learning. In this work, we present a universal parameter-efficient transfer learn…

2023

Accelerating Vision-Language Pretraining With Free Language Modeling

CVPR 2023poster

The state of the arts in vision-language pretraining (VLP) achieves exemplary performance but suffers from high training costs resulting from slow convergence and long training time, especially on large-scale web datasets. An essential obstacle to training efficiency lies in the entangled prediction…

2023

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

CVPR 2023poster

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this…

2023

Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models

ICCV 2023poster

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great effectiveness in transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to overfit to seen classes while failing to generalize to unsee…

Cited by 37PDFScholar
2023

Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models

ICCV 2023oral

Vision-language pre-training (VLP) models have shown vulnerability to adversarial examples in multimodal tasks. Furthermore, malicious adversaries can be deliberately transferred to attack other black-box models. However, existing work has mainly focused on investigating white-box attacks. In this p…

Cited by 65PDFcodeScholar
2023

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

ICCV 2023poster

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language models (LLMs) has made significant progress. However, we have observed and empirically demonstrated that these methods a…

Cited by 53PDFcodeScholar
2022

VLMixer: Unpaired Vision-Language Pre-training via Cross-Modal CutMix

ICML 2022spotlight

Existing vision-language pre-training (VLP) methods primarily rely on paired image-text datasets, which are either annotated by enormous human labors or crawled from the internet followed by elaborate data cleaning techniques. To reduce the dependency on well-aligned image-text pairs, it is promisin…

2021

End-to-End Dense Video Captioning With Parallel Decoding

ICCV 2021poster

Dense video captioning aims to generate multiple associated captions with their temporal locations from the video. Previous methods follow a sophisticated "localize-then-describe" scheme, which heavily relies on numerous hand-crafted components. In this paper, we proposed a simple yet effective fram…

Cited by 238PDFcodeScholar
2021

Hybrid Adaptive Control Strategy for Continuum Surgical Robot Under External Load

RA-L 2021

Natural orifice transluminal endoscopic surgery (NOTES) has received significant attentions due to its minimal incision trauma compared with traditional multi-port robot assisted surgery. Continuum robot can be used in NOTES due to its high flexibility which can adapt to circuitous paths. However, t

Cited by 45SourceScholar