← Search

Zhongyu Ouyang

12 accepted papers

2025

Growing Through Experience: Scaling Episodic Grounding in Language Models

ACL 2025long

Language models (LMs) require effective episodic grounding—the ability to learn from and apply past experiences—to perform well at physical planning tasks. While current approaches struggle with scalability and integration of episodic memory, which is particularly limited for medium-sized LMs (7B pa…

Cited by 0SourcePDFScholar
2025

Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making

EMNLP 2025

Modern embodied AI uses multimodal large language models (MLLMs) as policy models, predicting actions from final-layer hidden states. This widely adopted approach, however, assumes that monolithic last-layer representations suffice for decision-making—a structural simplification at odds with decades

Cited by 0SourcePDFScholar
2025

Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner

ICML 2025spotlight

Theory-of-mind (ToM) enables humans to infer mental states—such as beliefs, desires, and intentions—forming the foundation of social cognition. Existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning but struggle with scalability in multimod…

Cited by 0SourcePDFScholar
2025

Pretrained Image-Text Models are Secretly Video Captioners

NAACL 2025short

Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational resources and without complex modifications to address vide…

2025

SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models

EMNLP 2025

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic approach that involves a capable base model, high-quality rea

2025

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

NAACL 2025findings

Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended tempora…

2025

Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification

ACL 2025finding

Indigenous languages remain largely invisible in commercial language identification (LID) systems, a stark reality exemplified by Google Translate’s LangID tool, which supports over 100 languages but excludes all 150 Indigenous languages of North America. This technological marginalization is partic…

Cited by 0SourcePDFScholar
2024

From Coarse to Fine: Enable Comprehensive Graph Self-supervised Learning with Multi-granular Semantic Ensemble

ICML 2024oral

Self-supervised learning (SSL) has gained increasing attention in the graph learning community, owing to its capability of enabling powerful models pre-trained on large unlabeled graphs for general purposes, facilitating quick adaptation to specific domains. Though promising, existing graph SSL fram…

Cited by 2SourcePDFScholar
2024

GCVR: Reconstruction from Cross-View Enable Sufficient and Robust Graph Contrastive Learning

UAI 2024poster

Among the existing self-supervised learning (SSL) methods for graphs, graph contrastive learning (GCL) frameworks usually automatically generate supervision by transforming the same graph into different views through graph augmentation operations. The computation-efficient augmentation techniques e…

2024

Learning Musical Representations for Music Performance Question Answering

EMNLP 2024finding

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the audio-video QA demonstrate impressive capabilities on genera…

2024

Working Memory Identifies Reasoning Limits in Language Models

EMNLP 2024main

This study explores the inherent limitations of large language models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks—a benchmark used to assess wo…

Cited by 8SourcePDFScholar
2023

When Sparsity Meets Contrastive Models: Less Graph Data Can Bring Better Class-Balanced Representations

ICML 2023poster

Graph Neural Networks (GNNs) are powerful models for non-Euclidean data, but their training is often accentuated by massive unnecessary computation: on the one hand, training on non-Euclidean data has relatively high computational cost due to its irregular density properties; on the other hand, the…

Cited by 13SourcePDFScholar