← Search

Chunhui Zhang

28 accepted papers

2026

Mind the Gap: The Divergence Between Human and LLM-Generated Tasks

AAAI 2026technical

Humans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-g

Cited by 0SourcePDFScholar
2026

Uncertainty-Constrained Trustworthiness for Graph Learning

ICML 2026poster

Graph learning has been increasingly deployed in critical and sensitive domains, raising pressing demands for trustworthiness-robustness, fairness, and beyond. However, these properties are often undermined by various perturbations, which induce distributional uncertainty and compromise the trustwor…

Cited by 0SourceScholar
2025

Growing Through Experience: Scaling Episodic Grounding in Language Models

ACL 2025long

Language models (LMs) require effective episodic grounding—the ability to learn from and apply past experiences—to perform well at physical planning tasks. While current approaches struggle with scalability and integration of episodic memory, which is particularly limited for medium-sized LMs (7B pa…

Cited by 0SourcePDFScholar
2025

Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages

NAACL 2025short

Endangered languages, such as Navajo—the most widely spoken Native American language—are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This study evaluates Google’s Language Identification (LangID) tool, wh…

Cited by 2SourcePDFScholar
2025

Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making

EMNLP 2025

Modern embodied AI uses multimodal large language models (MLLMs) as policy models, predicting actions from final-layer hidden states. This widely adopted approach, however, assumes that monolithic last-layer representations suffice for decision-making—a structural simplification at odds with decades

Cited by 0SourcePDFScholar
2025

Learning Sparsity for Effective and Efficient Music Performance Question Answering

ACL 2025short

Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reasoning. Recent Music Performance Audio-Visual Question Answering (Music AVQA) datasets have been proposed to reflect these…

Cited by 0SourcePDFScholar
2025

MambaTrack: Exploiting Dual-Enhancement for Night UAV Tracking

ICASSP 2025accepted

Night unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, l…

Cited by 0SourceScholar
2025

Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models

ACL 2025long

Generative models such as Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) trained on massive datasets can lead them to memorize and inadvertently reveal sensitive information, raising ethical and privacy concerns. While some prior works have explored this issue in the conte…

2025

Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner

ICML 2025spotlight

Theory-of-mind (ToM) enables humans to infer mental states—such as beliefs, desires, and intentions—forming the foundation of social cognition. Existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning but struggle with scalability in multimod…

Cited by 0SourcePDFScholar
2025

Pretrained Image-Text Models are Secretly Video Captioners

NAACL 2025short

Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational resources and without complex modifications to address vide…

2025

SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models

EMNLP 2025

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic approach that involves a capable base model, high-quality rea

2025

Superficial Self-Improved Reasoners Benefit from Model Merging

EMNLP 2025

Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as such data becomes increasingly scarce, model self-improvement offers a promising alternative. However, this process can lead to model collapse, as the model’s output becomes overly deterministic with reduced diversity. I

2025

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

NAACL 2025findings

Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended tempora…

2025

Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification

ACL 2025finding

Indigenous languages remain largely invisible in commercial language identification (LID) systems, a stark reality exemplified by Google Translate’s LangID tool, which supports over 100 languages but excludes all 150 Indigenous languages of North America. This technological marginalization is partic…

Cited by 0SourcePDFScholar
2024

Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction

ACL 2024long

We introduce EVLGen, a streamlined framework designed for the pre-training of visually conditioned language generation models with high computational demands, utilizing frozen pre-trained large language models (LLMs). The conventional approach in vision-language pre-training (VLP) typically involves…

2024

GCVR: Reconstruction from Cross-View Enable Sufficient and Robust Graph Contrastive Learning

UAI 2024poster

Among the existing self-supervised learning (SSL) methods for graphs, graph contrastive learning (GCL) frameworks usually automatically generate supervision by transforming the same graph into different views through graph augmentation operations. The computation-efficient augmentation techniques e…

2024

Learning Musical Representations for Music Performance Question Answering

EMNLP 2024finding

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the audio-video QA demonstrate impressive capabilities on genera…

2024

Mitigating Emergent Robustness Degradation while Scaling Graph Learning

ICLR 2024poster

Although graph neural networks have exhibited remarkable performance in various graph tasks, a significant concern is their vulnerability to adversarial attacks. Consequently, many defense methods have been proposed to alleviate the deleterious effects of adversarial attacks and learn robust graph r…

Cited by 12SourcePDFScholar
2024

WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark

NeurIPS 2024poster

Underwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take t…

2024

Working Memory Identifies Reasoning Limits in Language Models

EMNLP 2024main

This study explores the inherent limitations of large language models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks—a benchmark used to assess wo…

Cited by 8SourcePDFScholar
2023

Boosting Graph Neural Networks via Adaptive Knowledge Distillation

AAAI 2023technical

Graph neural networks (GNNs) have shown remarkable performance on diverse graph mining tasks. While sharing the same message passing framework, our study shows that different GNNs learn distinct knowledge from the same graph. This implies potential performance improvement by distilling the complemen…

Cited by 42SourcePDFScholar
2023

Chasing All-Round Graph Representation Robustness: Model, Training, and Optimization

ICLR 2023poster

Graph Neural Networks (GNNs) have achieved state-of-the-art results on a variety of graph learning tasks, however, it has been demonstrated that they are vulnerable to adversarial attacks, raising serious security concerns. A lot of studies have been developed to train GNNs in a noisy environment an…

Cited by 21SourcePDFScholar
2023

Heterogeneous Graph Masked Autoencoders

AAAI 2023technical

Generative self-supervised learning (SSL), especially masked autoencoders, has become one of the most exciting learning paradigms and has shown great potential in handling graph data. However, real-world graphs are always heterogeneous, which poses three critical challenges that existing methods ign…

2023

When Sparsity Meets Contrastive Models: Less Graph Data Can Bring Better Class-Balanced Representations

ICML 2023poster

Graph Neural Networks (GNNs) are powerful models for non-Euclidean data, but their training is often accentuated by massive unnecessary computation: on the one hand, training on non-Euclidean data has relatively high computational cost due to its irregular density properties; on the other hand, the…

Cited by 13SourcePDFScholar
2022

Co-Modality Graph Contrastive Learning for Imbalanced Node Classification

NeurIPS 2022accept

Graph contrastive learning (GCL), leveraging graph augmentations to convert graphs into different views and further train graph neural networks (GNNs), has achieved considerable success on graph benchmark datasets. Yet, there are still some gaps in directly applying existing GCL methods to real-worl…

2022

Label-invariant Augmentation for Semi-Supervised Graph Classification

NeurIPS 2022accept

Recently, contrastiveness-based augmentation surges a new climax in the computer vision domain, where some operations, including rotation, crop, and flip, combined with dedicated algorithms, dramatically increase the model generalization and robustness. Following this trend, some pioneering attempts…

Cited by 39SourcePDFScholar