← Search

Zhicheng Zhang

30 accepted papers

2026

COMI: Coarse-to-fine Context Compression via Marginal Information Gain

ICLR 2026poster

Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse tasks. However, their deployment in long context scenarios remains hindered by computational inefficiency and information redundancy. Context compression methods address these challenges by significantly reducing…

Cited by 0SourcecodeScholar
2026

CaT-Diff: Cascaded Text-enhanced Diffusion Model for Time-Series Imputation

AAAI 2026technical

Most state-of-the-art time series imputation methods can leverage textual information to improve imputation quality, but they often struggle because they fail to effectively filter noisy information from large language model (LLM) derived textual information. Some existing solutions only filter over

Cited by 0SourcePDFScholar
2026

CollectiveKV: Decoupling and Sharing Collaborative Information in Sequential Recommendation

ICLR 2026poster

Sequential recommendation models are widely used in applications, yet they face stringent latency requirements. Mainstream models leverage the Transformer attention mechanism to improve performance, but its computational complexity grows with the sequence length, leading to a latency challenge for…

Cited by 0SourceScholar
2026

Learning Molecular Chirality via Chiral Determinant Kernels

ICLR 2026poster

Chirality is a fundamental molecular property that governs stereospecific behavior in chemistry and biology. Capturing chirality in machine learning models remains challenging due to the geometric complexity of stereochemical relationships and the limitations of traditional molecular representations…

Cited by 0SourcecodeScholar
2026

Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR Prediction

AAAI 2026technical

User behavior sequences in modern recommendation systems exhibit significant length heterogeneity, ranging from sparse short-term interactions to rich long-term histories. While longer sequences provide more context, we observe that increasing the maximum input sequence length in existing CTR models

Cited by 0SourcePDFScholar
2026

QPrompt-R1: Real-Time Reasoning for Domain-Generalized Semantic Segmentation via Group-Relative Query Alignment

ICLR 2026poster

Deploying semantic segmentation in driving and robotics requires both real-time inference and robustness to domain shifts, formalized as Real-Time Domain-Generalized Semantic Segmentation (RT-DGSS), which has not been fully addressed. Existing methods often treat real-time(RT) inference and domain g…

Cited by 0SourceScholar
2026

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus

ICLR 2026poster

Spine disorders affect 619 million people globally and are a leading cause of disability, yet AI-assisted diagnosis remains limited by the lack of level-aware, multimodal datasets. Clinical decision-making for spine disorders requires sophisticated reasoning across X-ray, CT, and MRI at specific ver…

Cited by 0SourceScholar
2025

AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose Estimation

IROS 2025

Recent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominan

Cited by 0SourceScholar
2025

BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold Networks

ICASSP 2025accepted

In this paper, we propose an efficient and drum-attuned beat tracking method based on Kolmogorov-Arnold networks (KAN). Traditional MLP-based frameworks struggle with complex musical signals due to limited capacity in modeling intricate patterns. Inspired by KAN’s efficient ability to capture comple…

Cited by 0SourceScholar
2025

Category-aware EEG Image Generation Based on Wavelet Transform and Contrast Semantic Loss

IJCAI 2025

Reconstructing visual stimuli from EEG signals is a crucial step in realizing brain-computer interfaces. In this paper, we propose a transformer-based EEG signal encoder integrating the Discrete Wavelet Transform (DWT) and the gating mechanism. Guided by the feature alignment and category-aware fusi

2025

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

NAACL 2025long

Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This challenge has prompted research on enhancing LLM-codebase interaction at a repository scale. Current solutions rely on similarity-based retrieval or manual…

2025

Cross-MoE: An Efficient Temporal Prediction Framework Integrating Textual Modality

EMNLP 2025

It has been demonstrated that incorporating external information as textual modality can effectively improve time series forecasting accuracy. However, current multi-modal models ignore the dynamic and different relations between time series patterns and textual features, which leads to poor perform

2025

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

NAACL 2025findings

Since the release of ChatGPT, large language models (LLMs) have demonstrated remarkable capabilities across various domains. A key challenge in developing these general capabilities is efficiently sourcing diverse, high-quality data. This becomes especially critical in reasoning-related tasks with s…

Cited by 2SourcePDFScholar
2025

GeMIMO: Searching the Cores of X-formers for Time Series Forecasting

ICASSP 2025accepted

In recent years, Transformer-based models have been widely used in time series forecasting tasks, demonstrating exceptional performance. However, these models lack interpretability, making it difficult to identify which components play a core role in predictions and which are redundant. To address t…

Cited by 0SourceScholar
2025

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

ICML 2025spotlight

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, p…

Cited by 0SourcePDFScholar
2025

MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection

ICASSP 2025accepted

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch a…

Cited by 0SourceScholar
2025

Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition

AAAI 2025technical

Given the critical role of birds in ecosystems, Fine-Grained Bird Recognition (FGBR) has gained increasing attention, particularly in distinguishing birds within similar subcategories. Although Vision Transformer (ViT)-based methods often outperform Convolutional Neural Network (CNN)-based methods i…

Cited by 0SourcePDFScholar
2025

M³HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

ICML 2025poster

Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordinated environments. We introduce Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality ($\…

Cited by 0SourcePDFScholar
2025

Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios

NAACL 2025findings

Large language models (LLMs) demonstrate exceptional capabilities in various scenarios. However, they suffer from much redundant information and are sensitive to the position of key information in long context scenarios. To address these challenges, we present Perception Compressor, a training-free…

Cited by 1SourcePDFScholar
2025

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

NeurIPS 2025poster

Understanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open…

Cited by 0SourceScholar
2024

A Framework for Inference Inspired by Human Memory Mechanisms

ICLR 2024poster

How humans and machines make sense of current inputs for relation reasoning and question-answering while putting the perceived information into context of our past memories, has been a challenging conundrum in cognitive science and artificial intelligence. Inspired by human brain's memory system and…

2024

Breast Ultrasound Computer-Aided Diagnosis Using Structure-Aware Triplet Path Networks

ICASSP 2024accepted

Breast ultrasound (BUS) is an effective imaging modality for breast cancer diagnosis. The structural characteristics of breast lesions play an important role in computer-aided diagnosis. In this paper, a novel structure-aware triplet path network (SATPN) was designed to integrate classification and…

Cited by 0SourceScholar
2024

ExtDM: Distribution Extrapolation Diffusion Model for Video Prediction

CVPR 2024poster

Video prediction is a challenging task due to its nature of uncertainty especially for forecasting a long period. To model the temporal dynamics advanced methods benefit from the recent success of diffusion models and repeatedly refine the predicted future frames with 3D spatiotemporal U-Net. Howeve…

Cited by 22SourcePDFScholar
2024

LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented Diffusion

CVPR 2024poster

Camouflaged vision perception is an important vision task with numerous practical applications. Due to the expensive collection and labeling costs this community struggles with a major bottleneck that the species category of its datasets is limited to a small number of object species. However the ex…

2024

MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution Distillation

CVPR 2024poster

Limited training data is a long-standing problem for video emotion analysis (VEA). Existing works leverage the power of large-scale image datasets for transferring while failing to extract the temporal correlation of affective cues in the video. Inspired by psychology research and empirical theory w…

Cited by 9SourcePDFScholar
2024

SRECT: Machine-Specific Spatial-Resolution Enhancement in Computed Tomography

ICASSP 2024accepted

Computed Tomography (CT) is an advanced imaging technology. To obtain high-resolution (HR) CT images from low-resolution (LR) sinograms, we present a deep-learning (DL) based CT super-resolution (SR) method.The proposed method combines a SR model in the sinogram domain and the iterative framework in…

Cited by 0SourceScholar
2023

Protein Representation Learning via Knowledge Enhanced Primary Structure Reasoning

ICLR 2023poster

Protein representation learning has primarily benefited from the remarkable development of language models (LMs). Accordingly, pre-trained protein models also suffer from a problem in LMs: a lack of factual knowledge. The recent solution models the relationships between protein and associated knowle…

Cited by 24SourcePDFScholar
2023

Weakly Supervised Video Emotion Detection and Prediction via Cross-Modal Temporal Erasing Network

CVPR 2023poster

Automatically predicting the emotions of user-generated videos (UGVs) receives increasing interest recently. However, existing methods mainly focus on a few key visual frames, which may limit their capacity to encode the context that depicts the intended emotions. To tackle that, in this paper, we p…

2021

MapGo: Model-Assisted Policy Optimization for Goal-Oriented Tasks

IJCAI 2021poster

In Goal-oriented Reinforcement learning, relabeling the raw goals in past experience to provide agents with hindsight ability is a major solution to the reward sparsity problem. In this paper, to enhance the diversity of relabeled goals, we develop FGI (Foresight Goal Inference), a new relabeling st…