← Search

Zeyuan Yang

7 accepted papers

2026

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

CVPR 2026

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render explicit images, but the heavy image-generation pre-training oft

Cited by 0SourcecodeScholar
2025

Rethinking Long Context Generation from the Continual Learning Perspective

COLING 2025main

Due to the limited context window, Large Language Models (LLMs) struggle with processing long contexts. Although fine-tuning can extend the context window, it incurs substantial computation costs. In contrast, recent tuning-free approaches reallocate the attention mechanism or incorporate temporary…

Cited by 1SourcePDFScholar
2025

VCA: Video Curious Agent for Long Video Understanding

ICCV 2025poster

Long video understanding poses unique challenges due to its temporal complexity and low information density. Recent works address this task by sampling numerous frames or incorporating auxiliary tools using LLMs, both of which result in high computational costs. In this work, we introduce a curiosit…

Cited by 0SourcePDFScholar
2024

Position: Towards Unified Alignment Between Agents, Humans, and Environment

ICML 2024poster

The rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, re…

Cited by 4SourcePDFScholar
2024

RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation

CVPR 2024poster

We leverage Large Language Models (LLM) for zeroshot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for reinforcement learning yet achieve relatively low success rates and lack generalizability. The intermittent nature of auditory signals further…

Cited by 7SourcePDFScholar
2024

UBSoft: A Simulation Platform for Robotic Skill Learning in Unbounded Soft Environments

CoRL 2024poster

It is desired to equip robots with the capability of interacting with various soft materials as they are ubiquitous in the real world. While physics simulations are one of the predominant methods for data collection and robot training, simulating soft materials presents considerable challenges. Spec…

Cited by 1SourcecodeScholar
2023

Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule Accumulation

EMNLP 2023long main

Large Language Models (LLMs) have showcased impressive performance. However, due to their inability to capture relationships among samples, these frozen LLMs inevitably keep repeating similar mistakes. In this work, we propose our Tuning-free Rule Accumulation (TRAN) framework, which guides LLMs in…

Cited by 0SourcecodeScholar