← Search

Changwen Zheng

31 accepted papers

2026

COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs

CVPR 2026

Despite Multimodal Large Language Models (MLLMs) having shown impressive capabilities, they may suffer from hallucinations. Empirically, we find that MLLMs attend disproportionately to task-irrelevant background regions compared with text-only LLMs, implying spurious background-answer correlations.

Cited by 0SourceScholar
2026

Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models

AAAI 2026technical

Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstre

Cited by 0SourcePDFScholar
2026

Exploring Transferability of Self-Supervised Learning by Task Conflict Calibration

AAAI 2026technical

In this paper, we explore the transferability of SSL by addressing two central questions: (i) what is the representation transferability of SSL, and (ii) how can we effectively model this transferability? Transferability is defined as the ability of a representation learned from one task to support

Cited by 0SourcePDFScholar
2026

Group Causal Policy Optimization for Post-Training Large Language Models

AAAI 2026technical

Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative reward

Cited by 0SourcePDFScholar
2026

HTG-GCL: Leveraging Hierarchical Topological Granularity from Cellular Complexes for Graph Contrastive Learning

AAAI 2026technical

Graph contrastive learning (GCL) aims to learn discriminative semantic invariance by contrasting different views of the same graph that share critical topological patterns. However, existing GCL approaches with structural augmentations often struggle to identify task-relevant topological structures,

Cited by 0SourcePDFScholar
2026

On the Plasticity and Stability for Post-Training Large Language Models

ICML 2026poster

Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability retention. We identify a root cause as the geometric conflict between plasticity and stability gradients, which leads t…

Cited by 0SourceScholar
2026

Supporting Multimodal Intermediate Fusion with Informatic Constraint and Distribution Coherence

ICLR 2026poster

Based on the prevalent intermediate fusion (IF) and late fusion (LF) frameworks, multimodal representation learning (MML) demonstrates its superiority over unimodal representation learning. To investigate the intrinsic factors underlying the empirical success of MML, research grounded in theoretical…

Cited by 0SourceScholar
2026

TMAE:Learning Targeted Multi-Agent Exploration via Causal Inference

AAAI 2026technical

Exploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the

Cited by 0SourcePDFScholar
2025

Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach

AAAI 2025technical

Graph representation learning methods are highly effective in handling complex non-Euclidean data by capturing intricate relationships and features within graph structures. However, traditional methods face challenges when dealing with heterogeneous graphs that contain various types of nodes and edg…

2025

LLM Enhancers for GNNs: An Analysis from the Perspective of Causal Mechanism Identification

ICML 2025poster

The use of large language models (LLMs) as feature enhancers to optimize node representations, which are then used as inputs for graph neural networks (GNNs), has shown significant potential in graph representation learning. However, the fundamental properties of this approach remain underexplored.…

Cited by 0SourcePDFScholar
2025

Learn to Think: Bootstrapping LLM Logic Through Graph Representation Learning

IJCAI 2025

Large Language Models (LLMs) have achieved remarkable success across various domains. However, they still face significant challenges, including high computational costs for training and limitations in solving complex reasoning problems. Although existing methods have extended the reasoning capabili

2025

Learning Invariant Causal Mechanism from Vision-Language Models

ICML 2025poster

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and…

Cited by 0SourcePDFScholar
2025

Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs

NeurIPS 2025poster

Large language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectiveness and efficiency, often encouraging unnecessarily long reasoning chains and wasting tokens. To address this, we propose…

Cited by 0SourceScholar
2025

Less Yet Robust: Crucial Region Selection for Scene Recognition

ICASSP 2025accepted

Scene recognition, particularly for aerial and underwater images, often suffers from various types of degradation, such as blurring or overexposure. Previous works that focus on convolutional neural networks have been shown to be able to extract panoramic semantic features and perform well on scene…

Cited by 0SourceScholar
2025

MAP: Supporting Multimodal Knowledge Graph Completion via Augmented Modality Alignment and Instance Preserving

ICASSP 2025accepted

Multimodal knowledge graphs (KGs) have found widespread applications in data integration and processing, yet existing multimodal knowledge graphs are often highly incomplete, which impedes their wide adoption. Thereby multimodal knowledge graph completion (MKGC) has attracted widespread attention. H…

Cited by 0SourceScholar
2025

On the Out-of-Distribution Generalization of Self-Supervised Learning

ICML 2025poster

In this paper, we focus on the out-of-distribution (OOD) generalization of self-supervised learning (SSL). By analyzing the mini-batch construction during the SSL training phase, we first give one plausible explanation for SSL having OOD generalization. Then, from the perspective of data generation…

2025

Towards the Causal Complete Cause of Multi-Modal Representation Learning

ICML 2025poster

Multi-Modal Learning (MML) aims to learn effective representations across modalities for accurate predictions. Existing methods typically focus on modality consistency and specificity to learn effective representations. However, from a causal perspective, they may lead to representations that contai…

Cited by 0SourcePDFScholar
2024

BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction

ICLR 2024poster

As a novel and effective fine-tuning paradigm based on large-scale pre-trained language models (PLMs), prompt-tuning aims to reduce the gap between downstream tasks and pre-training objectives. While prompt-tuning has yielded continuous advancements in various tasks, such an approach still remains a…

2024

Hacking Task Confounder in Meta-Learning

IJCAI 2024poster

Meta-learning enables rapid generalization to new tasks by learning knowledge from various tasks. It is intuitively assumed that as the training progresses, a model will acquire richer knowledge, leading to better generalization performance. However, our experiments reveal an unexpected result: ther…

2024

Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive Learning

AAAI 2024technical

Graph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discriminative unsupervised graph representation learning approach, GCL achieves impres…

2024

Radardiff: Improving Sea Clutter Suppression Using Diffusion Models for Radar Images

ICASSP 2024accepted

Marine radar is employed across multiple fields, notably in navigation, meteorology, defense, and security. Marine radar images are highly sensitive to sea clutter, highlighting the crucial importance of sea clutter suppression in radar image processing. However, existing algorithms for sea clutter…

Cited by 0SourceScholar
2024

Rethinking Causal Relationships Learning in Graph Neural Networks

AAAI 2024technical

Graph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite…

2024

Rethinking Dimensional Rationale in Graph Contrastive Learning from Causal Perspective

AAAI 2024technical

Graph contrastive learning is a general learning paradigm excelling at capturing invariant information from diverse perturbations in graphs. Recent works focus on exploring the structural rationale from graphs, thereby increasing the discriminability of the invariant information. However, such metho…

2024

T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration

AAAI 2024technical

Communication stands as a potent mechanism to harmonize the behaviors of multiple agents. However, existing work primarily concentrates on broadcast communication, which not only lacks practicality, but also leads to information redundancy. This surplus, one-fits-all information could adversely impa…

Cited by 10SourcePDFScholar
2023

Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal Perspective

AAAI 2023technical

Few-shot learning models learn representations with limited human annotations, and such a learning paradigm demonstrates practicability in various tasks, e.g., image classification, object detection, etc. However, few-shot object detection methods suffer from an intrinsic defect that the limited tra…

2023

Robust Causal Graph Representation Learning against Confounding Effects

AAAI 2023technical

The prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-prune…

2022

Bootstrapping Informative Graph Augmentation via A Meta Learning Approach

IJCAI 2022poster

Recent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented grap…

2022

Interventional Contrastive Learning with Meta Semantic Regularizer

ICML 2022spotlight

Contrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tes…

Cited by 34SourcePDFScholar
2022

MetAug: Contrastive Learning via Meta Feature Augmentation

ICML 2022spotlight

What matters for contrastive learning? We argue that contrastive learning heavily relies on informative features, or “hard” (positive or negative) features. Early works include more informative features by applying complex data augmentations and large batch size or memory bank, and recent works desi…

Cited by 41SourcePDFScholar
2022

MetaMask: Revisiting Dimensional Confounder for Self-Supervised Learning

NeurIPS 2022accept

As a successful approach to self-supervised learning, contrastive learning aims to learn invariant information shared among distortions of the input sample. While contrastive learning has yielded continuous advancements in sampling strategy and architecture design, it still remains two persistent de…

Cited by 16SourcePDFScholar
2022

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

NeurIPS 2022accept

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential…