← Search

Jian Cao

30 accepted papers

2026

DFRec: Dual Fluctuation Modeling of Multi-level Intent Evolution for Next-Item Recommendation

AAAI 2026technical

User sequential behaviors are driven by a variety of complex and evolving intents. Capturing the dynamic change of user intents has become critical yet challenging in the next-item recommendation. Existing studies usually model the transition relationships among multiple intents within a session or

Cited by 0SourcePDFScholar
2026

Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and Mixture

AAAI 2026technical

Market making (MM) through Reinforcement Learning (RL) has attracted significant attention in financial trading. With the development of Large Language Models (LLMs), more and more attempts are being made to apply LLMs to financial areas. A simple, direct application of LLM as an agent shows signifi

Cited by 0SourcePDFScholar
2025

ADELA: Accelerating Evolutionary Design of Machine Learning Pipelines with the Accompanying Surrogate Model

AAAI 2025technical

The end-to-end automated design of machine learning (ML) pipelines significantly reduces the workload for data scientists and democratizes ML for non-experts. Evolutionary algorithm (EA)-based automated ML (AutoML) systems, a prominent category of AutoML, often face inefficiencies due to the costly…

2025

COLUR: Confidence-Oriented Learning, Unlearning and Relearning with Noisy-Label Data for Model Restoration and Refinement

IJCAI 2025

Large deep learning models have achieved significant success in various tasks. However, the performance of a model can significantly degrade if it is needed to train on datasets with noisy labels with misleading or ambiguous information. To date, there are limited investigations on how to restore pe

Cited by 0SourcePDFScholar
2025

DABL: Detecting Semantic Anomalies in Business Processes Using Large Language Models

AAAI 2025technical

Detecting anomalies in business processes is crucial for ensuring operational success. While many existing methods rely on statistical frequency to detect anomalies, it's important to note that infrequent behavior doesn't necessarily imply undesirability. To address this challenge, detecting anomali…

2025

Enhancing Graph-based Fraud Detection by Adversarial Confidence Reweighting

ICASSP 2025accepted

Graph-based fraud detection has emerged as a pivotal tool in risk management, leveraging the power of graph neural networks to enhance node representations by aggregating information from neighboring nodes. However, this aggregation process can sometimes introduce noise by incorporating neighbors fr…

Cited by 0SourceScholar
2025

EoT: Evolution of Thoughts for Complex Reasoning Tasks

EMNLP 2025

Knowledge-based complex reasoning remains a significant challenge for large language models (LLMs) with in-context learning. To tackle this issue, previous studies focus on ensuring behavior fidelity, factuality, or reliability in generated reasoning processes that guide LLMs to produce solutions. H

2025

KaeDe: Progressive Generation of Logical Forms via Knowledge-Aware Question Decomposition for Improved KBQA

EMNLP 2025

Knowledge base question answering (KBQA) refers to the task of answering natural language questions using large-scale structured knowledge bases (KBs). Existing semantic parsing-based (SP-based) methods achieve superior performance by directly converting questions into structured logical form (LF) q

2025

OmniTry: Virtual Try-On Anything without Masks

NeurIPS 2025poster

Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g., jewelries and accessories, with mask-free setting for more pract…

Cited by 0SourcecodeScholar
2025

PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs

ICML 2025spotlight

The rise of generative APIs has fueled interest in privacy-preserving synthetic data generation. While the Private Evolution (PE) algorithm generates Differential Privacy (DP) synthetic images using diffusion model APIs, it struggles with few-shot private data due to the limitations of its DP-protec…

2025

Promoting Knowledge Base Question Answering by Directing LLMs to Generate Task-relevant Logical Forms

AAAI 2025technical

Knowledge base question answering (KBQA) refers to the system that produces answers to user queries by reasoning with a large-scale structured knowledge base. Advanced works have achieved great success either by generating logical forms (LF) or directly generating answers. Although the former typica…

Cited by 0SourcePDFScholar
2025

Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy

IJCAI 2025

Machine Unlearning (MU) technology facilitates the removal of the influence of specific data instances from trained models on request. Despite rapid advancements in MU technology, its vulnerabilities are still underexplored, posing potential risks of privacy breaches through leaks of ostensibly unle

Cited by 0SourcePDFScholar
2024

A Facial Expression Transfer Method Based on 3DMM and Diffusion Models

ICASSP 2024accepted

Due to the complex geometric features of facial expressions, achieving realistic facial expression transfer is a challenging task. This paper proposes a two-stage facial expression transfer method based on 3D Morphable Model (3DMM) and diffusion models. In the first stage, 3DMM-based DECA model is e…

Cited by 0SourceScholar
2024

An Upload-Efficient Scheme for Transferring Knowledge From a Server-Side Pre-trained Generator to Clients in Heterogeneous Federated Learning

CVPR 2024poster

Heterogeneous Federated Learning (HtFL) enables collaborative learning on multiple clients with different model architectures while preserving privacy. Despite recent research progress knowledge sharing in HtFL is still difficult due to data and model heterogeneity. To tackle this issue we leverage…

2024

Boosting 3D Visual Grounding by Object-Centric Referring Network

IROS 2024poster

3D visual grounding is tasked with locating a specific object within a 3D scene, as described by a given textual reference. This task is challenging because it requires (1) the accurate recognition of various objects in a 3D scene and (2) the understanding of spatial relations in the description. Ho…

Cited by 0SourceScholar
2024

Boosting Pruned Networks with Linear Over-Parameterization

ICASSP 2024accepted

Structured pruning is a popular technique for reducing the computational cost and memory footprint of neural networks by removing channels. It often leads to a decrease in network accuracy, which can be restored through fine-tuning. However, as the pruning ratio increases, it becomes progressively m…

Cited by 0SourceScholar
2024

FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning

AAAI 2024technical

Recently, Heterogeneous Federated Learning (HtFL) has attracted attention due to its ability to support heterogeneous models and data. To reduce the high communication cost of transmitting model parameters, a major challenge in HtFL, prototype-based HtFL methods are proposed to solely share class re…

2024

J-MAE: Jigsaw Meets Masked Autoencoders in X-Ray Security Inspection

ICASSP 2024accepted

The X-ray security inspection aims to identify any restricted items to protect public safety. Due to the lack of focus on unsupervised learning in this field, using pre-trained models on natural images leads to suboptimal results in downstream tasks. Previous works would lose the relative positional…

Cited by 0SourceScholar
2024

R2V-MIF: Rule-to-Vector Contrastive Learning and Multi-channel Information Fusion for Therapy Recommendation

IJCAI 2024poster

Integrating data-driven and rule-based approaches is crucial for therapy recommendations since they can collaborate to achieve better performance. Medical rules, which are chains of reasoning that can infer therapies, widely exist. However, their symbolic and logical forms make integrating them with…

2024

SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language Model

ICASSP 2024accepted

Embodied intelligence based on vision-language models aims to learn from interactions and derive general intelligence. However, existing generalized vision-language models cannot understand domain knowledge in home scenarios due to the lack of sweeping robot multimodal datasets. In this paper, we pr…

Cited by 0SourceScholar
2023

A New ANN-SNN Conversion Method with High Accuracy, Low Latency and Good Robustness

IJCAI 2023poster

Due to the advantages of low energy consumption, high robustness and fast inference speed, Spiking Neural Networks (SNNs), with good biological interpretability and the potential to be applied on neuromorphic hardware, are regarded as the third generation of Artificial Neural Networks (ANNs). Despit…

Cited by 18SourcePDFScholar
2023

Eliminating Domain Bias for Federated Learning in Representation Space

NeurIPS 2023poster

Recently, federated learning (FL) is popular for its privacy-preserving and collaborative learning abilities. However, under statistically heterogeneous scenarios, we observe that biased data domains on clients cause a representation bias phenomenon and further degenerate generic representations dur…

2023

GPFL: Simultaneously Learning Global and Personalized Feature Information for Personalized Federated Learning

ICCV 2023poster

Federated Learning (FL) is popular for its privacy-preserving and collaborative learning capabilities. Recently, personalized FL (pFL) has received attention for its ability to address statistical heterogeneity and achieve personalization in FL. However, from the perspective of feature extraction, m…

Cited by 64PDFcodeScholar
2023

Masked Distillation with Receptive Tokens

ICLR 2023poster

Distilling from the feature maps can be fairly effective for dense prediction tasks since both the feature discriminability and localization information can be well transferred. However, not every pixel contributes equally to the performance, and a good student should learn from what really matters…

2023

RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement Learning

AAAI 2023technical

Whole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learnin…

2023

Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization

CVPR 2023poster

Towards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant rol…

Cited by 5SourcePDFScholar
2023

Variational Sparse Inverse Cholesky Approximation for Latent Gaussian Processes via Double Kullback-Leibler Minimization

ICML 2023poster

To achieve scalable and accurate inference for latent Gaussian processes, we propose a variational approximation based on a family of Gaussian distributions whose covariance matrices have sparse inverse Cholesky (SIC) factors. We combine this variational approximation of the posterior with a similar…

2022

A free lunch from ViT: adaptive attention multi-scale fusion Transformer for fine-grained visual recognition

ICASSP 2022accepted

Learning subtle representation about object parts plays a vital role in fine-grained visual recognition (FGVR) field. The vision transformer (ViT) achieves promising results on computer vision due to its attention mechanism. Nonetheless, with the fixed size of patches in ViT, the class token in deep…

Cited by 0SourceScholar
2022

RAPQ: Rescuing Accuracy for Power-of-Two Low-bit Post-training Quantization

IJCAI 2022poster

We introduce a Power-of-Two post-training quantization( PTQ) method for deep neural network that meets hardware requirements and does not call for long-time retraining. PTQ requires a small set of calibration data and is easier for deployment, but results in lower accuracy than Quantization-Aware Tr…