← Search

zijian zhang

43 accepted papers

2026

$AutoDrive\text{-}P^3$: Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning

ICLR 2026poster

Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer from two major limitations: 1) Some VLMs directly output planning results without…

Cited by 0SourceScholar
2026

Beyond Illumination: Fine-Grained Detail Preservation in Extreme Dark Image Restoration

AAAI 2026technical

Recovering fine-grained details in extremely dark images remains challenging due to severe structural information loss and noise corruption. Existing enhancement methods often fail to preserve intricate details and sharp edges, limiting their effectiveness in downstream applications like text and ed

Cited by 0SourcePDFScholar
2026

Boosting Fine-Grained Urban Flow Inference via Lightweight Architecture and Focalized Optimization

AAAI 2026technical

Fine-grained urban flow inference is crucial for urban planning and intelligent transportation systems, enabling precise traffic management and resource allocation. However, the practical deployment of existing methods is hindered by two key challenges: the prohibitive computational cost of over-par

Cited by 0SourcePDFScholar
2026

Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction

AAAI 2026technical

Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on implicit network modeling of spatiotemporal relationships and motion characteristics, falling into the passive learning tr

Cited by 0SourcePDFScholar
2026

FINSENTLLM: MULTI-LLM AND STRUCTURED SEMANTIC SIGNALS FOR ENHANCED FINANCIAL SENTIMENT FORECASTING

ICASSP 2026poster

Financial sentiment analysis (FSA) has attracted significant attention, and recent studies increasingly explore large language models (LLMs) for this field. Yet most work evaluates only classification metrics, leaving unclear whether sentiment signals align with market behavior. We propose FinSentLL…

Cited by 0SourcePDFScholar
2026

FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification

IJCAI 2026

Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. However, existing approaches generally struggle to reconcile fine-grained representation learning, especially under c

Cited by 0Scholar
2026

GRAPE: Generalizing Robot Policy Via Preference Alignment

ICRA 2026poster

Despite the recent advancements of vision-language-action (VLA) models on a variety of robotics tasks, they suffer from critical issues such as poor generalizability to unseen tasks, due to their reliance on behavior cloning exclusively from successful rollouts. Furthermore, they are typically fine-…

2026

Generative Branching for Mixed-Integer Linear Programming

AAAI 2026technical

Branch-and-bound (B&B) is a fundamental algorithmic framework for solving Mixed-Integer Linear Programming (MILP) problems, where branching decisions critically affect solver efficiency. Recent learning-based methods apply imitation learning to select branching variables, but their deterministic pre

Cited by 0SourcePDFScholar
2026

HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting

AAAI 2026technical

Accurate traffic forecasting plays a vital role in intelligent transportation systems, enabling applications such as congestion control, route planning, and urban mobility optimization. However, traffic forecasting remains challenging due to two key factors: (1) complex spatial dependencies arising

Cited by 0SourcePDFScholar
2026

Minimum-Length Conformal Prediction Sets for Ordinal Classification

AAAI 2026technical

Ordinal classification has been widely applied in many high-stakes applications, e.g., medical imaging and diagnosis, where reliable uncertainty quantification (UQ) is essential for decision making. Conformal prediction (CP) is a general UQ framework that provides statistically valid guarantees, whi

Cited by 0SourcePDFScholar
2026

Reading the Cell, Designing the Cure: Perturbation-Conditioned Molecular Diffusion for Function-Oriented Drug Design

ICML 2026poster

When reliable target structures are unavailable at scale or phenotypes arise from dysregulated pathways, transcriptomic perturbations provide a system-level functional readout for drug action. In this work, we formalize Transcriptome-based Drug Design (TBDD) as a generative inverse problem: designin…

Cited by 0SourceScholar
2026

SPJFNet: Self-Mining Prior-Guided Joint Frequency Enhancement for Ultra-Efficient Dark Image Restoration

AAAI 2026technical

Current dark image restoration methods suffer from severe efficiency bottlenecks, primarily stemming from: computational burden and error correction costs associated with reliance on external priors (manual or cross-modal); redundant operations in complex multi-stage enhancement pipelines; and indis

Cited by 0SourcePDFScholar
2026

StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning

ICML 2026poster

Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to dependencies on both GPU kernel efficiency and host-side settings. Although LLM-based methods show promise on automated GPU kernel generation, prior works mainly foc…

Cited by 0SourceScholar
2025

Anyprefer: An Agentic Framework for Preference Data Synthesis

ICLR 2025poster

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its…

Cited by 0SourcePDFScholar
2025

CWNet: Causal Wavelet Network for Low-Light Image Enhancement

ICCV 2025poster

Traditional Low-Light Image Enhancement (LLIE) methods primarily focus on uniform brightness adjustment, often neglecting instance-level semantic information and the inherent characteristics of different features. To address these limitations, we propose CWNet (Causal Wavelet Network), a novel archi…

2025

Co-training with Progressive Distribution Alignment and Uncertainty-Interactive Relabeling for Semi-Supervised Domain Adaptive Semantic Segmentation

ICASSP 2025accepted

Self-training is a strong baseline for semi-supervised domain adaptive semantic segmentation. However, it inevitably introduces biased links between features and concepts in the prediction of certain "hard pixels", which may mislead the generalization of models. We consider these hard pixels to come…

Cited by 0SourceScholar
2025

GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle Dispatching

AAAI 2025technical

As urban residents demand higher travel quality, vehicle dispatch has become a critical component of online ride-hailing services. However, current vehicle dispatch systems struggle to navigate the complexities of urban traffic dynamics, including unpredictable traffic conditions, diverse driver beh…

Cited by 0SourcePDFScholar
2025

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

NeurIPS 2025poster

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modular…

Cited by 0SourcecodeScholar
2025

LLM-Powered User Simulator for Recommender System

AAAI 2025technical

User simulators can rapidly generate a large volume of timely user behavior data, providing a testing platform for reinforcement learning-based recommender systems, thus accelerating their iteration and optimization. However, prevalent user simulators generally suffer from significant limitations, i…

2025

Multifaceted User Modeling in Recommendation: A Federated Foundation Models Approach

AAAI 2025technical

Multifaceted user modeling aims to uncover fine-grained patterns and learn representations from user data, revealing their diverse interests and characteristics, such as profile, preference, and personality. Recent studies on foundation model-based recommendation have emphasized the Transformer arch…

2025

OCLNet: Obfuscation feature Contrastive Learning Network for Weakly Supervised Semantic Segmentation on Ultrasound Images

ICASSP 2025accepted

Deep learning-based semantic segmentation technology has become a critical tool in assisting doctors with automatic lesion segmentation in medical images. However, the high cost of acquiring large-scale, pixel-level annotations poses a significant challenge, limiting the scalability and application…

Cited by 0SourceScholar
2025

Pancreatic Cancer Diagnosis System

AAAI 2025technical

My research direction is about the pancreatic cancer diagnosis system. Pancreatic cancer, as one of the cancers with the highest mortality rate, has always been a difficult problem in world medicine. I hope that through my efforts, I can contribute to the integration of AI and medicine, and contribu…

Cited by 0SourcePDFScholar
2025

S-RAG: A Novel Audit Framework for Detecting Unauthorized Use of Personal Data in RAG Systems

ACL 2025long

Retrieval-Augmented Generation (RAG) systems combine external data retrieval with text generation and have become essential in applications requiring accurate and context-specific responses. However, their reliance on external data raises critical concerns about unauthorized collection and usage of…

2025

Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations

NeurIPS 2025poster

Generative models have recently gained attention in recommendation systems by directly predicting item identifiers from user interaction sequences. However, existing methods suffer from significant information loss due to the separation of stages such as quantization and sequence modeling, hindering…

Cited by 0SourceScholar
2025

TimeEmb: A Lightweight Static-Dynamic Disentanglement Framework for Time Series Forecasting

NeurIPS 2025poster

Temporal non-stationarity, the phenomenon that time series distributions change over time, poses fundamental challenges to reliable time series forecasting. Intuitively, the complex time series can be decomposed into two factors, i.e., time-invariant and time-varying components, which indicate stati…

Cited by 0SourcecodeScholar
2024

A Soft Robotic Gripper With a Belt Loop Actuated Adhesion Design for Gentle Handling of Fragile Object

RA-L 2024

This study presents a soft gripper incorporating a novel belt loop actuated adhesion design, offering a comprehensive solution for the fabrication, variable-scale actuation, and contact sensing of directional adhesives. The solution facilitates high-resolution control of directional adhesives with r

Cited by 5SourceScholar
2024

Federated Adaptation for Foundation Model-based Recommendations

IJCAI 2024poster

With the recent success of large language models, particularly foundation models with generalization abilities, applying foundation models for recommendations becomes a new paradigm to improve existing recommendation systems. It becomes a new open challenge to enable the foundation model to capture…

2024

Image Understanding Makes for A Good Tokenizer for Image Generation

NeurIPS 2024poster

Modern image generation (IG) models have been shown to capture rich semantics valuable for image understanding (IU) tasks. However, the potential of IU models to improve IG performance remains uncharted. We address this issue using a token-based IG framework, which relies on effective tokenizers to…

2024

LLM-ESR: Large Language Models Enhancement for Long-tailed Sequential Recommendation

NeurIPS 2024spotlight

Sequential recommender systems (SRS) aim to predict users' subsequent choices based on their historical interactions and have found applications in diverse fields such as e-commerce and social media. However, in real-world systems, most users interact with only a handful of items, while the majority…

Cited by 10SourcePDFScholar
2024

Multiple-Source Localization from a Single-Snapshot Observation Using Graph Bayesian Optimization

AAAI 2024technical

Due to the significance of its various applications, source localization has garnered considerable attention as one of the most important means to confront diffusion hazards. Multi-source localization from a single-snapshot observation is especially relevant due to its prevalence. However, the inher…

2023

AutoSTL: Automated Spatio-Temporal Multi-Task Learning

AAAI 2023technical

Spatio-temporal prediction plays a critical role in smart city construction. Jointly modeling multiple spatio-temporal tasks can further promote an intelligent city life by integrating their inseparable relationship. However, existing studies fail to address this joint learning problem well, which g…

Cited by 28SourcePDFScholar
2023

Dual Personalization on Federated Recommendation

IJCAI 2023poster

Federated recommendation is a new Internet service architecture that aims to provide privacy-preserving recommendation services in federated settings. Existing solutions are used to combine distributed recommendation algorithms and privacy-preserving mechanisms. Thus it inherently takes the form of…

2023

MSDC: Exploiting Multi-State Power Consumption in Non-intrusive Load Monitoring Based on a Dual-CNN Model

AAAI 2023technical

Non-intrusive load monitoring (NILM) aims to decompose aggregated electrical usage signal into appliance-specific power consumption and it amounts to a classical example of blind source separation tasks. Leveraging recent progress on deep learning techniques, we design a new neural NILM model {\em M…

2023

ShiftDDPMs: Exploring Conditional Diffusion Models by Shifting Diffusion Trajectories

AAAI 2023technical

Diffusion models have recently exhibited remarkable abilities to synthesize striking image samples since the introduction of denoising diffusion probabilistic models (DDPMs). Their key idea is to disrupt images into noise through a fixed forward process and learn its reverse process to generate samp…

Cited by 16SourcePDFScholar
2022

An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization

ICASSP 2022accepted

Sound source separation and localization for situational awareness enables a wide range of applications such as hearing enhancement and audio beam-forming. We present an end-to-end deep learning framework to separate and localize multiple audio sources from the mixture of multi-channels. The propose…

Cited by 0SourceScholar
2022

Unsupervised Representation Learning from Pre-trained Diffusion Probabilistic Models

NeurIPS 2022accept

Diffusion Probabilistic Models (DPMs) have shown a powerful capacity of generating high-quality image samples. Recently, diffusion autoencoders (Diff-AE) have been proposed to explore DPMs for representation learning via autoencoding. Their key idea is to jointly train an encoder for discovering mea…

2021

Distributed Dynamic Map Fusion via Federated Learning for Intelligent Networked Vehicles

ICRA 2021poster

The technology of dynamic map fusion among networked vehicles has been developed to enlarge sensing ranges and improve sensing accuracies for individual vehicles. This paper proposes a federated learning (FL) based dynamic map fusion framework to achieve high map quality despite unknown numbers of o…

Cited by 87SourcecodeScholar
2021

From Local to Global Norm Emergence: Dissolving Self-reinforcing Substructures with Incremental Social Instruments

ICML 2021spotlight

Norm emergence is a process where agents in a multi-agent system establish self-enforcing conformity through repeated interactions. When such interactions are confined to a social topology, several self-reinforcing substructures (SRS) may emerge within the population. This prevents a formation of a…

Cited by 12SourcePDFScholar
2019

REM: From Structural Entropy to Community Structure Deception

NeurIPS 2019poster

This paper focuses on the privacy risks of disclosing the community structure in an online social network. By exploiting the community affiliations of user accounts, an attacker may infer sensitive user attributes. This raises the problem of community structure deception (CSD), which asks for ways t…