← Search

Yunlong Tang

12 accepted papers

2026

Gaze-Based Teleoperation with Intent Inference Model for Robotic Manipulators

ICRA 2026poster

Eye gaze-based control interfaces provide a non-invasive means of enhancing human-robot collaboration for activities of daily living and can reduce the cognitive burden on operators performing complex tasks. Eye gaze has traditionally been used for "gaze triggering," where fixating on an object acti…

Cited by 0Scholar
2026

When to Think and When to Look: Uncertainty-Guided Lookback

CVPR 2026

Test-time "thinking" (i.e., generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision-language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how t

Cited by 0SourcecodeScholar
2025

CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion

AAAI 2025technical

Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like memory and cognition. Among these top-down influences, language plays a crucial role in guiding attention by shaping how vi…

Cited by 7SourcePDFScholar
2025

Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding

AAAI 2025technical

Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets, their temporal understanding capacity in video-language tas…

Cited by 5SourcePDFScholar
2025

Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability

NeurIPS 2025poster

Vision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. Many previous studies have observed that adversarial examples crafted on ViTs exhibit higher transferability than those crafted on CNNs, indicating that ViTs c…

Cited by 0SourceScholar
2025

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

NeurIPS 2025poster

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits: text-to-image (T2I) benchmarks that lacks multi-modal condition…

Cited by 0SourceScholar
2025

MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

NeurIPS 2025poster

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understandin…

Cited by 0SourcecodeScholar
2025

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

CVPR 2025poster

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on linguistic data, effectively interpret and process visual content? Th…

Cited by 4SourcePDFScholar
2025

V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning

AAAI 2025technical

Video summarization aims to create short, accurate, and cohesive summaries of longer videos. Despite the existence of various video summarization datasets, a notable limitation is their limited amount of source videos, which hampers the effective training of advanced large vision-language models (VL…

Cited by 69SourcePDFScholar
2025

VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?

CVPR 2025poster

The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multimodal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessmen…

2025

ZeroSep: Separate Anything in Audio with Zero Training

NeurIPS 2025poster

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the…

Cited by 0SourceScholar
2025

p-AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?

ICCV 2025poster

The Audio-Visual Acoustic Synthesis (AVAS) task aims to model realistic audio propagation behavior within a specific visual scene. Prior works often rely on sparse image representations to guide acoustic synthesis. However, we argue that this approach is insufficient to capture the intricate physica…

Cited by 0SourcePDFScholar