← Search

Yaolun Zhang

4 accepted papers

2026

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding

CVPR 2026

We revisit video hallucination in multimodal large language models (Video-MLLMs) from a semantic aggregation perspective. While prior work attributes hallucinations to language priors, missing frames, or visual encoder biases, these explanations overlook errors arising during the aggregation of corr

Cited by 0SourcecodeScholar
2026

EVA: Efficient Reinforcement Learning for End-to-End Video Agent

CVPR 2026

Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames.Existing approaches typically treat MLLMs as passive recognizers, processing entire videos or uniformly

Cited by 0SourcecodeScholar
2026

Position: Digital Agents Require Unified Agent-Native Environments

ICML 2026poster

Large language models (LLMs) are increasingly deployed as digital agents that perform multi-step digital work on a computer, but the environments in which they operate remain fragmented and task-specific. Our position is that digital agents need Agent-Native Computer: interfaces that expose system c…

Cited by 0SourceScholar
2025

MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines

ICML 2025poster

Large Language Models (LLMs) have demonstrated the ability to solve a wide range of practical tasks within multi-agent systems. However, existing human-designed multi-agent frameworks are typically limited to a small set of pre-defined scenarios, while current automated design methods suffer from se…

Cited by 0SourcePDFScholar