← Search

Longbing Cao

30 accepted papers

2026

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

ICML 2026poster

Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the continuous evolution of modeling dynamics, existing VAR token reduction methods face three key limitations: heuristic stage partition, non-adaptive sch…

Cited by 0SourceScholar
2026

CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records

AAAI 2026technical

Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inabili

Cited by 0SourcePDFScholar
2026

Markovian Scale Prediction: A New Era of Visual Autoregressive Generation

CVPR 2026

Visual AutoRegressive modeling (VAR) based on next-scale prediction has revitalized autoregressive visual generation. Although its full-context dependency, i.e., modeling all previous scales for next-scale prediction, facilitates more stable and comprehensive representation learning by leveraging co

Cited by 0SourceScholar
2026

PatchET: Learning Enzyme Temperature Properties Through Patch-Based Neural Architectures

AAAI 2026technical

Understanding enzyme thermal properties is essential for biotechnology and protein engineering, yet experimental measurements of attributes such as temperature optimum, stability, and range remain labor-intensive and costly. Prior studies have shown that specific regions within enzyme sequences disp

Cited by 0SourcePDFScholar
2026

Swordsman: Entropy-Driven Adaptive Block Partition for Efficient Diffusion Language Models

ICML 2026poster

Block-wise decoding effectively improves the inference speed and quality in diffusion language models (DLMs) by combining inter-block sequential denoising and intra-block parallel unmasking. However, existing block-wise decoding methods typically partition blocks in a rigid and fixed manner, which i…

Cited by 0SourceScholar
2026

VividFace: Real-Time and Realistic Facial Expression Shadowing for Humanoid Robots

ICRA 2026poster

Humanoid facial expression shadowing enables robots to realistically imitate human facial expressions in real time, which is critical for lifelike, facially expressive humanoid robots and affective human–robot interaction. Existing progress in humanoid facial expression imitation remains limited, of…

2025

Dynamic Spectral Graph Anomaly Detection

AAAI 2025technical

Graph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distrib…

2025

Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning

NeurIPS 2025poster

Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or cause…

Cited by 0SourceScholar
2025

Mixture of Online and Offline Experts for Non-Stationary Time Series

AAAI 2025technical

We consider a general and realistic scenario involving non-stationary time series, consisting of several offline intervals with different distributions within a fixed offline time horizon, and an online interval that continuously receives new samples. For non-stationary time series, the data distrib…

2025

Revealing Multimodal Causality with Large Language Models

NeurIPS 2025poster

Uncovering cause-and-effect mechanisms from data is fundamental to scientific progress. While large language models (LLMs) show promise for enhancing causal discovery (CD) from unstructured data, their application to the increasingly prevalent multimodal setting remains a critical challenge. Even wi…

Cited by 0SourcecodeScholar
2025

SCoT: Unifying Consistency Models and Rectified Flows via Straight-Consistent Trajectories

NeurIPS 2025poster

Pre-trained diffusion models are commonly used to generate clean data (e.g., images) from random noises, effectively forming pairs of noises and corresponding clean images. Distillation on these pre-trained models can be viewed as the process of constructing advanced trajectories within the pair to…

Cited by 0SourceScholar
2025

UGotMe: An Embodied System for Affective Human-Robot Interaction

ICRA 2025

Equipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affective human-robot interaction. However, enabling current vision-aware multimodal emotion recognition models for affective

Cited by 6SourcecodeScholar
2024

Frequency Spectrum Is More Effective for Multimodal Representation and Fusion: A Multimodal Spectrum Rumor Detector

AAAI 2024technical

Multimodal content, such as mixing text with images, presents significant challenges to rumor detection in social media. Existing multimodal rumor detection has focused on mixing tokens among spatial and sequential locations for unimodal representation or fusing clues of rumor veracity across modali…

2024

Rethinking Fourier Transform from A Basis Functions Perspective for Long-term Time Series Forecasting

NeurIPS 2024poster

The interaction between Fourier transform and deep learning opens new avenues for long-term time series forecasting (LTSF). We propose a new perspective to reconsider the Fourier transform from a basis functions perspective. Specifically, the real and imaginary parts of the frequency components can…

2024

Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled Data

NeurIPS 2024poster

There are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorit…

Cited by 0SourcePDFScholar
2023

FourierGNN: Rethinking Multivariate Time Series Forecasting from a Pure Graph Perspective

NeurIPS 2023poster

Multivariate time series (MTS) forecasting has shown great importance in numerous industries. Current state-of-the-art graph neural network (GNN)-based forecasting methods usually require both graph networks (e.g., GCN) and temporal networks (e.g., LSTM) to capture inter-series (spatial) dynamics an…

2023

Frequency-domain MLPs are More Effective Learners in Time Series Forecasting

NeurIPS 2023poster

Time series forecasting has played the key role in different industrial, including finance, traffic, energy, and healthcare domains. While existing literatures have designed many sophisticated architectures based on RNNs, GNNs, or Transformers, another kind of approaches based on multi-layer percept…

2022

A Probabilistic Code Balance Constraint with Compactness and Informativeness Enhancement for Deep Supervised Hashing

IJCAI 2022poster

Building on deep representation learning, deep supervised hashing has achieved promising performance in tasks like similarity retrieval. However, conventional code balance constraints (i.e., bit balance and bit uncorrelation) imposed on avoiding overfitting and improving hash code quality are unsuit…

2021

Coupling Macro-Sector-Micro Financial Indicators for Learning Stock Representations with Less Uncertainty

AAAI 2021technical

While the stock movement prediction has been intensively studied, existing work suffers from weak generalization because of the uncertainty in both data and modeling. On one hand, training a stock representation on stochastic stock data in an end-to-end manner may lead to excessive modeling, which i…

2021

Enlivening Redundant Heads in Multi-head Self-attention for Machine Translation

EMNLP 2021main

Multi-head self-attention recently attracts enormous interest owing to its specialized functions, significant parallelizable computation, and flexible extensibility. However, very recent empirical studies show that some self-attention heads make little contribution and can be pruned as redundant hea…

2021

Graph Learning based Recommender Systems: A Review

IJCAI 2021poster

Recent years have witnessed the fast development of the emerging topic of Graph Learning based Recommender Systems (GLRS). GLRS mainly employ advanced graph learning approaches to model users’ preferences and intentions as well as items’ characteristics and popularity for Recommender Systems (RS). D…

2021

Self-supervised Bilingual Syntactic Alignment for Neural Machine Translation

AAAI 2021technical

While various neural machine translation (NMT) methods have integrated mono-lingual syntax knowledge into the linguistic representation of sequence-to-sequence, no research is available on aligning the syntactic structures of target language with the corresponding source language syntactic structure…

2021

Tripartite Collaborative Filtering with Observability and Selection for Debiasing Rating Estimation on Missing-Not-at-Random Data

AAAI 2021technical

Most collaborative filtering (CF) models estimate missing ratings with an implicit assumption that the ratings are missing-at-random, which may cause the biased rating estimation and degraded performance since recent deep exploration shows that ratings may likely be missing-not-at-random (MNAR). To…

Cited by 14SourcePDFScholar
2020

Intention2Basket: A Neural Intention-driven Approach for Dynamic Next-basket Planning

IJCAI 2020poster

User purchase behaviours are complex and dynamic, which are usually observed as multiple choice actions across a sequence of shopping baskets. Most of the existing next-basket prediction approaches model user actions as homogeneous sequence data without considering complex and heteroge…

Cited by 0SourcePDFScholar