← Search

Zeren Zhang

8 accepted papers

2026

Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling

AAAI 2026technical

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV)

Cited by 0SourcePDFScholar
2025

AIR: Unifying Individual and Collective Exploration in Cooperative Multi-Agent Reinforcement Learning

AAAI 2025technical

Exploration in cooperative multi-agent reinforcement learning (MARL) remains challenging for value-based agents due to the absence of an explicit policy. Existing approaches include individual exploration based on uncertainty towards the system and collective exploration through behavioral diversity…

2025

Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver

ICASSP 2025accepted

Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems, which require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geo…

Cited by 0SourceScholar
2025

Enhancing Large Language Models on Domain-specific Tasks: A Novel Training Strategy via Domain Adaptation and Preference Alignment

ICASSP 2025accepted

In handling complex, domain-specific tasks, particularly in the context of state-owned assets and enterprises (SOAEs), general LLMs suffer from the knowledge gap due to insufficient exposure to domain-specific corpora, and the value disagreement, as they are aligned with universal values rather than…

Cited by 0SourceScholar
2025

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

ICASSP 2025accepted

Combining face-swapping with lip synchronization offers a cost-effective solution for generating customized talking faces. However, directly cascading existing models can introduce significant interference and reduce video clarity due to limited interaction space in the low-level RGB domain. To solv…

Cited by 0SourceScholar
2023

Consensus Learning for Cooperative Multi-Agent Reinforcement Learning

AAAI 2023technical

Almost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and ch…

Cited by 17SourcePDFScholar
2023

Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARL

NeurIPS 2023poster

Value decomposition methods have gained popularity in the field of cooperative multi-agent reinforcement learning. However, almost all existing methods follow the principle of Individual Global Max (IGM) or its variants, which limits their problem-solving capabilities. To address this, we propose a…

Cited by 4SourcePDFScholar
2023

PCSalmix: Gradient Saliency-Based Mix Augmentation for Point Cloud Classification

ICASSP 2023accepted

Point cloud classification has sparked many researchers’ interest for its cornerstone role in 3D applications. Inheriting the CutMix series augmentation that performs well in 2D images, PointCutMix and RSMix are proposed to generate new samples for 3D point clouds, by replacing partial points of one…

Cited by 0SourceScholar