← Search

Zhi Zhang

35 accepted papers

2026

EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing

ICLR 2026poster

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images—resulting in limited coverage and inheriting biases from prior generative models—or (ii) rely *solely* on zero-shot vis…

Cited by 0SourcecodeScholar
2026

Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing

ICLR 2026poster

Recently, visualization-of-thought (VoT) has unlocked new opportunities for complex spatial reasoning in multimodal large language models (MLLMs) by complementing verbal reasoning with visual thinking. However, the autoregressive accumulation of lengthy and redundant tokens substantially increases c…

Cited by 0SourceScholar
2026

ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns

ICML 2026poster

Mixture-of-Experts (MoE) effectively scales model capacity while preserving computational efficiency through sparse expert activation. However, training high-quality MoEs from scratch is prohibitively expensive. A promising alternative is to convert pretrained dense models into sparse MoEs. Existing…

Cited by 0SourceScholar
2026

From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression Comprehension

AAAI 2026technical

Recent advances in Referring Expression Comprehension (REC) have been largely driven by supervised learning on curated datasets, where each expression is assumed to refer to exactly one known object. However, such assumptions rarely hold in real-world scenarios, where expressions can refer to multip

Cited by 0SourcePDFScholar
2026

Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMs

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studi

Cited by 0SourcePDFScholar
2026

Music Atelier: Exploring the Knowing–Doing Gap in LLM Creativity via Symbolic Music Composition

IJCAI 2026

The knowing--doing gap, the mismatch between ideas articulated during model reasoning and the realized creative artifact, remains a fundamental challenge in creative AI and persists in LLM-based artistic creation. This paper responds to this gap by introducing a systematic and interpretable framewor

Cited by 0Scholar
2026

OmniScale: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

AAAI 2026technical

Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticat

Cited by 0SourcePDFScholar
2026

Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions

ICLR 2026poster

Generalized linear bandits have been extensively studied due to their broad applicability in real-world online decision-making problems. However, these methods typically assume that the expected reward function is known to the users, an assumption that is often unrealistic in practice. Misspecificat…

Cited by 0SourceScholar
2025

Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

NAACL 2025findings

Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasin…

Cited by 2SourcePDFScholar
2025

Cross-modal Information Flow in Multimodal Large Language Models

CVPR 2025poster

The recent advancements in auto-regressive multimodal large language models (MLLMs) have demonstrated promising progress for vision-language tasks. While there exists a variety of studies investigating the processing of linguistic information within large language models, little is currently known a…

2025

Efficient Utility-Preserving Machine Unlearning with Implicit Gradient Surgery

NeurIPS 2025poster

Machine unlearning (MU) aims to efficiently remove sensitive or harmful memory from a pre-trained model. The key challenge is to balance the potential tradeoff between unlearning efficacy and utility preservation, which involves forgetting undesirable information as defined while maintaining the mod…

Cited by 0SourcecodeScholar
2025

Let the Code LLM Edit Itself When You Edit the Code

ICLR 2025poster

In this work, we investigate a typical scenario in code generation where a developer edits existing code in real time and requests a code assistant, e.g., a large language model, to re-predict the next token or next line on the fly. Naively, the LLM needs to re-encode the entire KV cache to provide…

Cited by 0SourcePDFScholar
2025

Mixture of Knowledge Minigraph Agents for Literature Review Generation

AAAI 2025technical

Literature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel fram…

Cited by 0SourcePDFScholar
2025

NeuroAda: Activating Each Neuron’s Potential for Parameter-Efficient Fine-Tuning

EMNLP 2025

Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. The former, such as LoRA, introduce additional modules to adapt the model to downstream tasks, offering strong memory efficiency. However, their representation

2025

Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions

IROS 2025

The adaptivity and maneuvering capabilities of Autonomous Underwater Vehicles (AUVs) have drawn significant attention in oceanic research, due to the unpredictable disturbances and strong coupling among the AUV’s degrees of freedom. In this paper, we developed large language model (LLM)-enhanced rei

Cited by 10SourceScholar
2025

SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning

ICML 2025poster

Advances in Parameter-efficient Fine-tuning (PEFT) bridged the performance gap with Full Fine-Tuning (FFT) through sophisticated analysis of pre-trained parameter spaces. Starting from drawing insights from Neural Engrams (NE) in Biological Neural Networks (BNNs), we establish a connection between t…

2025

Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory

AISTATS 2025poster

Lifelong reinforcement learning (RL) has been developed as a paradigm for extending single-task RL to more realistic, dynamic settings. In lifelong RL, the "life" of an RL agent is modeled as a stream of tasks drawn from a task distribution. We propose EPIC (Empirical PAC-Bayes that Improves Continu…

Cited by 0SourceScholar
2024

Beyond Mimicking Under-Represented Emotions: Deep Data Augmentation with Emotional Subspace Constraints for EEG-Based Emotion Recognition

AAAI 2024technical

In recent years, using Electroencephalography (EEG) to recognize emotions has garnered considerable attention. Despite advancements, limited EEG data restricts its potential. Thus, Generative Adversarial Networks (GANs) are proposed to mimic the observed distributions and generate EEG data. However,…

Cited by 14SourcePDFScholar
2024

Formal Verification of Unknown Stochastic Systems via Non-parametric Estimation

AISTATS 2024poster

A novel data-driven method for formal verification is proposed to study complex systems operating in safety-critical domains. The proposed approach is able to formally verify discrete-time stochastic dynamical systems against temporal logic specifications only using observation samples and without t…

Cited by 3SourcePDFScholar
2024

Gradient-based Parameter Selection for Efficient Fine-Tuning

CVPR 2024poster

With the growing size of pre-trained models full fine-tuning and storing all the parameters for various downstream tasks is costly and infeasible. In this paper we propose a new parameter-efficient fine-tuning method Gradient-based Parameter Selection (GPS) demonstrating that only tuning a few selec…

2024

Multivariate Time Series Forecasting By Graph Attention Networks With Theoretical Guarantees

AISTATS 2024poster

Multivariate time series forecasting (MTSF) aims to predict future values of multiple variables based on past values of multivariate time series, and has been applied in fields including traffic flow prediction, stock price forecasting, and anomaly detection. Capturing the inter-dependencies among m…

Cited by 4SourcePDFScholar
2024

SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training

NeurIPS 2024poster

Recent years have witnessed a clear trend towards language models with an ever-increasing number of parameters, as well as the growing training overhead and memory usage. Distributed training, particularly through Sharded Data Parallelism (ShardedDP) which partitions optimizer states among workers,…

Cited by 2SourcePDFScholar
2024

Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation

ICML 2024poster

In this work, we leverage the intrinsic segmentation of language sequences and design a new positional encoding method called Bilevel Positional Encoding (BiPE). For each position, our BiPE blends an intra-segment encoding and an inter-segment encoding. The intra-segment encoding identifies the loca…

2023

GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Video

ICCV 2023poster

3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the perf…

Cited by 80PDFcodeScholar
2023

Uncertainty-Aware Few-Shot Class-Incremental Learning

ICASSP 2023accepted

In a real-world setting, machine needs to continuously recognize new categories without forgetting. However, the number of new categories may be small. For some difficult categories, even humans cannot recognize only based on few-shot examples. To address the above issues, an innovative uncertainty-…

Cited by 0SourceScholar
2022

Reinforcement Learning under a Multi-agent Predictive State Representation Model: Method and Theory

ICLR 2022spotlight

We study reinforcement learning for partially observable multi-agent systems where each agent only has access to its own observation and reward and aims to maximize its cumulative rewards. To handle partial observations, we propose graph-assisted predictive state representations (GAPSR), a scalable…

Cited by 12SourcePDFScholar
2021

CrossNorm and SelfNorm for Generalization Under Distribution Shifts

ICCV 2021poster

Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world applications, well-trained models with previous normalization…

Cited by 73PDFcodeScholar
2021

Progressive Coordinate Transforms for Monocular 3D Object Detection

NeurIPS 2021poster

Recognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While ther…

2020

TLPG-Tracker: Joint Learning of Target Localization and Proposal Generation for Visual Tracking

IJCAI 2020poster

Target localization and proposal generation are two essential subtasks in generic visual tracking, and it is a challenge to address both the two efficiently. In this paper, we propose an efficient two-stage architecture which makes full use of the complementarity of two subtasks to achieve robust lo…

Cited by 0SourcePDFScholar
2019

Bag of Tricks for Image Classification with Convolutional Neural Networks

CVPR 2019poster

Much of the recent progress made in image classification research can be credited to training procedure refinements, such as changes in data augmentations and optimization methods. In the literature, however, most refinements are either briefly mentioned as implementation details or only visible in…

Cited by 2019PDFcodeScholar
2017

Rate-coverage analysis and optimization for joint audio-video multimedia retrieval

ICASSP 2017accepted

In this work, we consider the problem of automatic content retrieval (ACR) using joint audio-video fingerprints. We focus on how to balance the query accuracy and the size of fingerprint, and how to allocate the fingerprint bits to video and audio frames to maximize the query accuracy. By introducin…

Cited by 0SourceScholar