← Search

Yidan Zhang

12 accepted papers

2026

FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding

CVPR 2026

Natural videos exhibit heterogeneous temporal dynamics, with certain segments undergoing high-dynamic scene transitions and others dominated by low-dynamic visual changes. However, treating all frames identically, a common practice in most MLLMs, leads to redundant visual encoding, which results in

Cited by 0SourcecodeScholar
2026

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities to remote sensing imagery remains challenging, as targets ar

Cited by 0SourcecodeScholar
2026

LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid

AAAI 2026technical

Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic level

Cited by 0SourcePDFScholar
2026

RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention

AAAI 2026technical

Video prediction is plagued by a fundamental trilemma: achieving high-resolution and perceptual quality typically comes at the cost of real-time speed, hindering its use in latency-critical applications. This challenge is most acute for autonomous UAVs in dense urban environments, where foreseeing e

Cited by 0SourcePDFScholar
2026

ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation

ICLR 2026poster

Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield severe reconstruction incompleteness. Recent advancements in diffusion-based 3D generative techniques offer the potential to…

Cited by 0SourcecodeScholar
2025

P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs

EMNLP 2025

Recent advancements in large language models (LLMs) showcase varied multilingual capabilities across tasks like translation, code generation, and reasoning. Previous assessments often limited their scope to fundamental natural language processing (NLP) or isolated capability-specific tasks. To allev

2025

Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders

ACL 2025long

The mechanisms behind multilingual capabilities in Large Language Models (LLMs) have been examined using neuron-based or internal-activation-based methods. However, these methods often face challenges such as superposition and layer-wise activation variance, which limit their reliability. Sparse Aut…

2024

Large Language Models Can Not Perform Well in Understanding and Manipulating Natural Language at Both Character and Word Levels?

EMNLP 2024finding

Despite their promising performance across various tasks, recent studies reveal that Large language models (LLMs) still exhibit significant deficiencies in handling several word-level and character-level tasks, e.g., word unscrambling and sentence editing, indicating urgent needs for substantial imp…

Cited by 2SourcePDFScholar
2024

Rationales for Answers to Simple Math Word Problems Confuse Large Language Models

ACL 2024findings

Recently, large language models (LLMs) have demonstrated breakthrough mathematical problem-solving capabilities in grade school math word problems (MWP). For example, on the MWP benchmark GSM8K, the accuracy of GPT-3.5-Turbo and MetaMath-70B reaches 80.80% and 82.30%, respectively. One question aris…

Cited by 0SourcePDFScholar
2023

Dynamic Voting for Efficient Reasoning in Large Language Models

EMNLP 2023long findings

Multi-path voting methods like Self-consistency have been used to mitigate reasoning errors in large language models caused by factual errors and illusion generation. However, these methods require excessive computing resources as they generate numerous reasoning paths for each problem. And our expe…

Cited by 0SourceScholar
2023

MVImgNet: A Large-Scale Dataset of Multi-View Images

CVPR 2023poster

Being data-driven is one of the most iconic properties of deep learning algorithms. The birth of ImageNet drives a remarkable trend of "learning from large-scale data" in computer vision. Pretraining on ImageNet to obtain rich universal representations has been manifested to benefit various 2D visua…

Cited by 180SourcePDFScholar
2023

Unifying Discrete and Continuous Representations for Unsupervised Paraphrase Generation

EMNLP 2023long main

Unsupervised paraphrase generation is a challenging task that benefits a variety of downstream NLP applications. Current unsupervised methods for paraphrase generation typically employ round-trip translation or denoising, which require translation corpus and result in paraphrases overly similar to t…

Cited by 0SourceScholar