← Search

Zhe LI

73 accepted papers

2026

Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate impressive cross-modal capabilities, yet their substantial size poses significant deployment challenges. Knowledge distillation (KD) is a promising solution for compressing these models, but existing methods primarily rely on static next-token ali…

Cited by 0SourceScholar
2026

Converge Faster, Talk Less: Hessian-Informed Federated Zeroth-Order Optimization

ICLR 2026poster

Zeroth-order (ZO) optimization enables dimension-free communication in federated learning (FL), making it attractive for fine-tuning of large language models (LLMs) due to significant communication savings. However, existing ZO-FL methods largely overlook curvature information, despite its well-esta…

Cited by 0SourceScholar
2026

DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations

CVPR 2026

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these diffusion models realizes high-fidelity disentangled control betw

Cited by 0SourceScholar
2026

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

CVPR 2026

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robots relies on explicit motion reconstruction, leading to cascaded errors, high lat

Cited by 0SourceScholar
2026

From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

ICLR 2026poster

Natural language offers a natural interface for humanoid robots, but existing text-to-motion pipelines remain cumbersome and unreliable. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone…

Cited by 0SourceScholar
2026

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models benefit from Chain-of-Thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We propose Latent Reasoning VLA (LaRA-VLA), a unified VLA framework…

Cited by 0SourceScholar
2026

Subgraph Encoding with Bicentric Sphere Node Labeling and Pooling for Link Prediction

AAAI 2026technical

Learning representation of the enclosing subgraph of node pairs is recognized as an efficient approach for link-oriented prediction tasks in network applications. The core challenge within this subgraph encoding approach is how to effectively distinguish and then properly aggregate the contribution

Cited by 0SourcePDFScholar
2026

ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

CVPR 2026

CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks.Existing MLLMs typically perform textual reasoning solely from a single static mathematical image, overlooking dynamic visual acquisition dur

Cited by 0SourcecodeScholar
2026

Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing

ICLR 2026poster

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their deployment is frequently undermined by undesirable behaviors such as generating harmful content, factual inaccuracies, and societal biases. Diagnosing the root causes of these failures poses a critical challenge for AI…

Cited by 0SourcecodeScholar
2025

Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization

ICLR 2025poster

Federated Learning (FL) offers a promising framework for collaborative and privacy-preserving machine learning across distributed data sources. However, the substantial communication costs associated with FL significantly challenge its efficiency. Specifically, in each communication round, the com…

2025

AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

CVPR 2025poster

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D mo…

2025

BodyGen: Advancing Towards Efficient Embodiment Co-Design

ICLR 2025spotlight

Embodiment co-design aims to optimize a robot's morphology and control policy simultaneously. While prior work has demonstrated its potential for generating environment-adaptive robots, this field still faces persistent challenges in optimization efficiency due to the (i) combinatorial nature of mo…

2025

Bridging the User-side Knowledge Gap in Knowledge-aware Recommendations with Large Language Models

AAAI 2025technical

In recent years, knowledge graphs have been integrated into recommender systems as item-side auxiliary information, enhancing recommendation accuracy. However, constructing and integrating structural user-side knowledge remains a significant challenge due to the improper granularity and inherent sca…

2025

DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings Learning

CVPR 2025poster

Vision-Language (VL) alignment across image and text modalities is a challenging task due to the inherent semantic ambiguity of data with multiple possible meanings. Existing methods typically solve it by learning multiple sub-representation spaces to encode each input data as a set of embeddings, a…

Cited by 0SourcePDFScholar
2025

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

ICCV 2025poster

Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scen…

2025

Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression

CVPR 2025poster

Leveraging the generative power of diffusion models, generative image compression has achieved impressive perceptual fidelity even at extremely low bitrates. However, current methods often neglect the non-uniform complexity of images, limiting their ability to balance global perceptual quality with…

Cited by 0SourcePDFScholar
2025

Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification

ICASSP 2025accepted

In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level…

Cited by 0SourceScholar
2025

Dual-Flow: Transferable Multi-Target, Instance-Agnostic Attacks via $\textit{In-the-wild}$ Cascading Flow Optimization

NeurIPS 2025poster

Adversarial attacks are widely used to evaluate model robustness, and in black-box scenarios, the transferability of these attacks becomes crucial. Existing generator-based attacks have excellent generalization and transferability due to their instance-agnostic nature. However, when training generat…

Cited by 0SourceScholar
2025

Exact and Linear Convergence for Federated Learning under Arbitrary Client Participation is Attainable

NeurIPS 2025poster

This work tackles the fundamental challenges in Federated Learning (FL) posed by arbitrary client participation and data heterogeneity, prevalent characteristics in practical FL settings. It is well-established that popular FedAvg-style algorithms struggle with exact convergence and can suffer from…

Cited by 0SourceScholar
2025

FAST: A Lightweight Mechanism Unleashing Arbitrary Client Participation in Federated Learning

IJCAI 2025

Federated Learning (FL) provides a flexible distributed platform where numerous clients with high data and system heterogeneity can collaborate to learn a model. While previous research has shown that FL can handle diverse data, it often completely assumes idealized conditions. In practice, real-wor

Cited by 0SourcePDFScholar
2025

HiRemate: Hierarchical Approach for Efficient Re-materialization of Neural Networks

ICML 2025poster

Training deep neural networks (DNNs) on memory-limited GPUs is challenging, as storing intermediate activations often exceeds available memory. Re-materialization, a technique that preserves exact computations, addresses this by selectively recomputing activations instead of storing them. However,…

Cited by 0SourcePDFScholar
2025

Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image Retrieval

ICCV 2025poster

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text describing the user's intention without training on the triplet datasets. The key to this task is to make specified changes to specific objects in the reference image based on the text…

Cited by 0SourcePDFScholar
2025

LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning

ICLR 2025poster

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP’s pretraining on static image-text pairs. This work introduces LaMP, a novel Lan…

2025

Plug-and-Play Multi-Domain Fusion Adaptation for Cross-Subject EEG-Based Motor Imagery Classification

ICRA 2025

Motor imagery (MI) classification in rehabilitation brain-computer interfaces (RBCIs) faces significant challenges due to the variability of electroencephalography (EEG) signals across subjects. Existing methods typically require extensive EEG data collection from each new subject, which is time-con

Cited by 1SourceScholar
2025

Spectral-Aware Low-Rank Adaptation for Speaker Verification

ICASSP 2025accepted

Previous research has shown that the principal singular vectors of a pre-trained model’s weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As a result, the LoRA-based parameter-efficient fine-tuning (P…

Cited by 0SourceScholar
2025

Subtractive Training for Music Stem Insertion Using Latent Diffusion Models

ICASSP 2025accepted

We present Subtractive Training<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, a simple and novel method for synthesizing individual musical instrument stems given other instruments as context. This method pairs a dataset of complete music mixe…

Cited by 0SourceScholar
2025

Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach

NeurIPS 2025poster

Split Federated Learning (SFL) enables scalable training on edge devices by combining the parallelism of Federated Learning (FL) with the computational offloading of Split Learning (SL). Despite its great success, SFL suffers significantly from the well-known straggler issue in distributed learning…

Cited by 0SourceScholar
2025

TrInk: Ink Generation with Transformer Network

EMNLP 2025

In this paper, we propose TrInk, a Transformer-based model for ink generation, which effectively captures global dependencies. To better facilitate the alignment between the input text and generated stroke points, we introduce scaled positional embeddings and a Gaussian memory mask in the cross-atte

2025

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

NeurIPS 2025spotlight

Constructing accurate digital twins of articulated objects is essential for robotic simulation training and embodied AI world model building, yet historically requires painstaking manual modeling or multi-stage pipelines. In this work, we propose \textbf{URDF-Anything}, an end-to-end automatic recon…

Cited by 0SourceScholar
2025

Utterance as A Bridge: Few-shot Joint Learning of Empathy Detection and Empathy Intent Classification

ICASSP 2025accepted

Empathy detection (ED) and empathy intent classification (EIC) aim to identify the empathy direction expressed in user utterances and the underlying empathy intent behind them. Previous studies show that facilitating information transfer between tasks can enhance model performance. However, the inte…

Cited by 0SourceScholar
2025

Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs

EMNLP 2025

Large Vision-Language Models (LVLMs) have made significant strides in multimodal comprehension, thanks to extensive pre-training and fine-tuning on large-scale visual datasets. However, despite their robust textual safety mechanisms, they remain vulnerable to harmful visual inputs. Existing safeguar

2024

Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling

CVPR 2024poster

Modeling animatable human avatars from RGB videos is a long-standing and challenging problem. Recent works usually adopt MLP-based neural radiance fields (NeRF) to represent 3D humans but it remains difficult for pure MLPs to regress pose-dependent garment details. To this end we introduce Animatabl…

2024

Are AI-Generated Text Detectors Robust to Adversarial Perturbations?

ACL 2024long

The widespread use of large language models (LLMs) has sparked concerns about the potential misuse of AI-generated text, as these models can produce content that closely resembles human-generated text. Current detectors for AI-generated text (AIGT) lack robustness against adversarial perturbations,…

2024

Capturing Detail Variations for Lightweight Neural Radiance Fields

ICASSP 2024accepted

Neural Radiance Fields (NeRF) has recently overhauled novel view synthesis, but it requires extensive computations for training and captures variations in detail with difficulty. In this paper, we propose a novel framework, termed CD-TDRF, to mitigate these dilemmas. CD-TDRF factorizes a density vox…

Cited by 0SourceScholar
2024

DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models

EMNLP 2024industry

Improving the efficiency of inference in Large Language Models (LLMs) is a critical area of research. Post-training Quantization (PTQ) is a popular technique, but it often faces challenges at low-bit levels, particularly in downstream tasks. Quantization-aware Training (QAT) can alleviate this probl…

Cited by 2SourcePDFScholar
2024

Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

EMNLP 2024finding

Large language models (LLMs) are increasingly being adopted in a wide range of real-world applications. Despite their impressive performance, recent studies have shown that LLMs are vulnerable to deliberately crafted adversarial prompts even when aligned via Reinforcement Learning from Human Feedbac…

2024

Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding

ICASSP 2024accepted

Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occur…

Cited by 0SourceScholar
2024

Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters

ICASSP 2024accepted

Fine-tuning a pre-trained Transformer model (PTM) for speech applications in a parameter-efficient manner offers the dual benefits of reducing memory and leveraging the rich feature representations in massive unlabeled datasets. However, existing parameter-efficient fine-tuning approaches either ada…

Cited by 0SourceScholar
2024

Efficiency Calibration of Implicit Regularization in Deep Networks via Self-paced Curriculum-Driven Singular Value Selection

IJCAI 2024poster

The generalization of neural networks has been a major focus of research in deep learning. It is often interpreted as an implicit bias towards solutions with specific properties. Especially, in practical applications, it has been observed that linear neural networks (LNN) tend to favor low-rank solu…

Cited by 0SourcePDFScholar
2024

Energy-based Backdoor Defense without Task-Specific Samples and Model Retraining

ICML 2024poster

Backdoor defense is crucial to ensure the safety and robustness of machine learning models when under attack. However, most existing methods specialize in either the detection or removal of backdoors, but seldom both. While few works have addressed both, these methods rely on strong assumptions or e…

Cited by 4SourcePDFScholar
2024

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

CVPR 2024poster

Creating high-fidelity 3D head avatars has always been a research hotspot but there remains a great challenge under lightweight sparse view setups. In this paper we propose Gaussian Head Avatar represented by controllable 3D Gaussians for high-fidelity head avatar modeling. We optimize the neutral 3…

2024

General Point Model Pretraining with Autoencoding and Autoregressive

CVPR 2024poster

The pre-training architectures of large language models encompass various types including autoencoding models autoregressive models and encoder-decoder models. We posit that any modality can potentially benefit from a large language model as long as it undergoes vector quantization to become discret…

2024

MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning

CVPR 2024poster

The scarcity of annotated data has sparked significant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medical visual representation learning. However existing research overlooks the multi-granularity nature of medical visual representation and la…

Cited by 15SourcePDFScholar
2024

MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos

ECCV 2024poster

"We present a novel pipeline for learning high-quality triangular human avatars from multi-view videos. Recent methods for avatar learning are typically based on neural radiance fields (NeRF), which is not compatible with traditional graphics pipeline and poses great challenges for operations like e…

2023

DialogMI: A Dialogue Model Based on Enhancing Dialogue Mutual Information

ICASSP 2023accepted

Most of the open-domain dialogue models tend to perform insufficiently in generating informative response. The possible reason is that they lack the capability of enhancing the mutual information between generated responses and dialogue history. To address this issue, we present a novel task of the…

Cited by 0SourceScholar
2023

Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space

ICASSP 2023accepted

The challenges in applying contrastive learning to speaker verification (SV) are that the softmax-based contrastive loss lacks discriminative power and that the hard negative pairs can easily influence learning. To overcome the first challenge, we propose a contrastive learning SV framework incorpor…

Cited by 0SourceScholar
2023

From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft Labels

ICCV 2023poster

Knowledge Distillation (KD) uses the teacher's prediction logits as soft labels to guide the student, while self-KD does not need a real teacher to require the soft labels. This work unifies the formulations of the two tasks by decomposing and reorganizing the generic KD loss into a Normalized KD (N…

Cited by 111PDFcodeScholar
2023

SMARTformer: Semi-Autoregressive Transformer with Efficient Integrated Window Attention for Long Time Series Forecasting

IJCAI 2023poster

The success of Transformers in long time series forecasting (LTSF) can be attributed to their attention mechanisms and non-autoregressive (NAR) decoder structures, which capture long-range de- pendencies. However, time series data also contain abundant local temporal dependencies, which are often ov…

Cited by 7SourcePDFScholar
2022

AvatarCap: Animatable Avatar Conditioned Monocular Human Volumetric Capture

ECCV 2022poster

"To address the ill-posed problem caused by partial observations in monocular human volumetric capture, we present AvatarCap, a novel framework that introduces animatable avatars into the capture pipeline for high-fidelity reconstruction in both visible and invisible regions. Our method firstly crea…

2022

Focal and Global Knowledge Distillation for Detectors

CVPR 2022poster

Knowledge distillation has been applied to image classification successfully. However, object detection is much more sophisticated and most knowledge distillation methods have failed on it. In this paper, we point out that in object detection, the features of the teacher and student vary greatly in…

Cited by 356PDFcodeScholar
2022

Masked Generative Distillation

ECCV 2022poster

"Knowledge distillation has been applied to various tasks successfully. The current distillation algorithm usually improves students’ performance by imitating the output of the teacher. This paper shows that teachers can also improve students’ representation power by guiding students’ feature recove…

2022

Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline

ECCV 2022poster

"Tracking in 3D scenes is gaining momentum because of its numerous applications in robotics, autonomous driving, and scene understanding. Currently, 3D tracking is limited to specific model-based approaches involving point clouds, which impedes 3D trackers from applying in natural 3D scenes. RGBD se…

2021

Implicit Feature Alignment: Learn To Convert Text Recognizer to Text Spotter

CVPR 2021poster

Text recognition is a popular research subject with many associated challenges. Despite the considerable progress made in recent years, the text recognition task itself is still constrained to solve the problem of reading cropped line text images and serves as a subtask of optical character recognit…

Cited by 16PDFcodeScholar
2021

Lightweight Multi-Person Total Motion Capture Using Sparse Multi-View Cameras

ICCV 2021poster

Multi-person total motion capture is extremely challenging when it comes to handle severe occlusions, different reconstruction granularities from body to face and hands, drastically changing observation scales and fast body movements. To overcome these challenges above, we contribute a lightweight t…

Cited by 64PDFScholar
2021

POSEFusion: Pose-Guided Selective Fusion for Single-View Human Volumetric Capture

CVPR 2021poster

We propose POse-guided SElective Fusion (POSEFusion), a single-view human volumetric capture method that leverages tracking-based methods and tracking-free inference to achieve high-fidelity and dynamic 3D reconstruction. By contributing a novel reconstruction framework which contains pose-guided ke…

Cited by 33PDFScholar
2019

A Robust Zero-Sum Game Framework for Pool-based Active Learning

AISTATS 2019poster

In this paper, we present a novel robust zero- sum game framework for pool-based active learning grounded on advanced statistical learning theory. Pool-based active learning usually consists of two components, namely, learning of a classifier given labeled data and querying of unlabeled data for lab…

Cited by 22SourcePDFScholar
2019

EIGEN: Ecologically-Inspired GENetic Approach for Neural Network Structure Searching From Scratch

CVPR 2019poster

Designing the structure of neural networks is considered one of the most challenging tasks in deep learning, especially when there is few prior knowledge about the task domain. In this paper, we propose an Ecologically-Inspired GENetic (EIGEN) approach that uses the concept of succession, extinction…

Cited by 33PDFScholar
2019

Learning from brains how to regularize machines

NeurIPS 2019poster

Despite impressive performance on numerous visual tasks, Convolutional Neural Networks (CNNs) --- unlike brains --- are often highly sensitive to small perturbations of their input, e.g. adversarial noise leading to erroneous decisions. We propose to regularize CNNs using large-scale neuroscience da…

Cited by 70SourcePDFScholar
2019

Prior-Aware Neural Network for Partially-Supervised Multi-Organ Segmentation

ICCV 2019accepted

Accurate multi-organ abdominal CT segmentation is essential to many clinical applications such as computer-aided intervention. As data annotation requires massive human labor from experienced radiologists, it is common that training data is usually partially-labeled. However, these background labels…

2018

A Simple Analysis for Exp-concave Empirical Minimization with Arbitrary Convex Regularizer

AISTATS 2018poster

In this paper, we present a simple analysis of fast rates with high probability of empirical minimization for it stochastic composite optimization over a finite-dimensional bounded convex set with exponential concave loss functions and an arbitrary convex regularization. To the best of our knowle…

Cited by 0SourcePDFScholar
2018

Adaptive Negative Curvature Descent with Applications in Non-convex Optimization

NeurIPS 2018poster

Negative curvature descent (NCD) method has been utilized to design deterministic or stochastic algorithms for non-convex optimization aiming at finding second-order stationary points or local minima. In existing studies, NCD needs to approximate the smallest eigen-value of the Hessian matrix with a…

Cited by 18SourcePDFScholar
2018

Thoracic Disease Identification and Localization With Limited Supervision

CVPR 2018poster

Accurate identification and localization of abnormalities from radiology images play an integral part in clinical diagnosis and treatment planning. Building a highly accurate prediction model for these tasks usually requires a large number of images manually annotated with labels and finding sites o…

Cited by 455SourcePDFScholar
2018

Widely Linear CLMS Based Cancelation of Nonlinear Self -Interference in Full-Duplex Direct-Conversion Transceivers

ICASSP 2018accepted

An augmented nonlinear complex LMS (ANCLMS) algorithm is proposed to adaptively mitigate both the linear and nonlinear self-interference (SI) components in a full-duplex direct-conversion transceiver (DCT). A data prewhitening scheme, which exploits the known SI signal distributions, is also adopted…

Cited by 3SourceScholar
2017

Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank

ICML 2017poster

Recently low displacement rank (LDR) matrices, or so-called structured matrices, have been proposed to compress large-scale neural networks. Empirical results have shown that neural networks with weight matrices of LDR matrices, referred as LDR neural networks, can achieve significant reduction in s…

Cited by 79SourcePDFScholar