← Search

Wenhui Tan

8 accepted papers

2026

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

ICLR 2026poster

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio, an effective benchmark must comprehensively cover three key a…

Cited by 0SourceScholar
2026

MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding

CVPR 2026

Efficiently understanding long-form videos remains a fundamental challenge for Multimodal Large Language Models (MLLMs). In this paper, we present MLLM-Sampler Joint Evolution (MSJoE), a novel framework that jointly evolves the MLLM and a lightweight key-frame sampler for efficient long-form video u

Cited by 0SourceScholar
2026

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

CVPR 2026

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitations. The fundamental reasons for this lie in two points: (1) long-form video und

Cited by 0SourceScholar
2026

Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

ICML 2026poster

Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases p…

Cited by 0SourceScholar
2026

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

ICML 2026poster

Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse reward signals and substantial computational overhead. We propose Video-OPD, an e…

Cited by 0SourceScholar
2025

Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains

NeurIPS 2025poster

Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient. In this paper, we introduce Compressed Latent Reasoning (CoLaR), a novel framework that dynamically compresses rea…

Cited by 0SourceScholar
2025

Think Then React: Towards Unconstrained Action-to-Reaction Motion Generation

ICLR 2025poster

Modeling human-like action-to-reaction generation has significant real-world applications, like human-robot interaction and games. Despite recent advancements in single-person motion generation, it is still challenging to well handle action-to-reaction generation, due to the difficulty of directly p…

Cited by 2SourcePDFScholar