← Search

KAI ZHAO

55 accepted papers

2026

Adaptive Coordinated Control of an Assistive Lower-Limb Exoskeleton for Hemiparetic Patients

RA-L 2026

Lower-limb exoskeletons play an important role in improving gait symmetry and walking ability for hemiparetic patients. However, individual variability in gait symmetry and motor capability limits their assistive performance. To address this challenge, this paper proposes an Adaptive Coordinated Con

Cited by 0SourceScholar
2026

Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision Recognition

AAAI 2026technical

Images are typically sampled on a uniform grid,despite their non-uniform information distribution—some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsamp

Cited by 0SourcePDFScholar
2026

CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards

AAAI 2026technical

Large-scale Chinese spelling correction (CSC) remains critical for real-world text processing, yet existing LLMs and supervised methods lack robustness to novel errors and rely on costly annotations. We introduce CEC-Zero, a zerosupervision reinforcement learning framework that addresses this by ena

Cited by 0SourcePDFScholar
2026

GeoDM: Geometry-aware Distribution Matching for Dataset Distillation

ICML 2026poster

Dataset distillation aims to synthesize a compact subset of the original data, enabling models trained on it to achieve performance comparable to those trained on the original large dataset. Existing distribution-matching methods are confined to Euclidean spaces, making them only capture linear stru…

Cited by 0SourceScholar
2026

INSTANCERSR: REAL-WORLD SUPER-RESOLUTION VIA INSTANCE-AWARE REPRESENTATION ALIGNMENT

ICASSP 2026poster

Existing real-world super-resolution (RSR) methods based on generative priors have achieved remarkable progress in producing high-quality and globally consistent reconstructions. However, they often struggle to recover fine-grained details of diverse object instances in complex real-world scenes. Th…

Cited by 0SourcePDFScholar
2026

MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction

ICML 2026poster

Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant,…

Cited by 0SourceScholar
2026

PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and Combination

AAAI 2026technical

In recent years, transformer-based models have achieved remarkable success in sensitive domains, including healthcare, finance and personalized services, but their deployment raises significant privacy concerns. Existing secure inference studies have introduced cryptographic techniques such as Homom

Cited by 0SourcePDFScholar
2026

SABER: Switchable and Balanced Training for Efficient LLM Reasoning

AAAI 2026technical

Large language models (LLMs) empowered by chain-of-thought reasoning have achieved impressive accuracy on complex tasks but suffer from excessive inference costs and latency when applied uniformly to all problems. We propose SABER (Switchable and Balanced Training for Efficient LLM Reasoning), a rei

Cited by 0SourcePDFScholar
2026

Temporal Dynamics Enhancer for Directly Trained Spiking Object Detectors

AAAI 2026technical

Spiking Neural Networks (SNNs), with their brain-inspired spatiotemporal dynamics and spike-driven computation, have emerged as promising energy-efficient alternatives to Artificial Neural Networks (ANNs). However, existing SNNs typically replicate inputs directly or aggregate them into frames at fi

Cited by 0SourcePDFScholar
2025

Cross-City Latent Space Alignment for Consistency Region Embedding

ICML 2025poster

Learning urban region embeddings has substantially advanced urban analysis, but their typical focus on individual cities leads to disparate embedding spaces, hindering cross-city knowledge transfer and the reuse of downstream task predictors. To tackle this issue, we present Consistent Region Embedd…

Cited by 0SourcePDFScholar
2025

Efficient Training of Neural Fractional-Order Differential Equation via Adjoint Backpropagation

AAAI 2025technical

Fractional-order differential equations (FDEs) enhance traditional differential equations by extending the order of differential operators from integers to real numbers, offering greater flexibility in modeling complex dynamic systems with nonlocal characteristics. Recent progress at the intersectio…

2025

Enhancing Diversity for Data-free Quantization

CVPR 2025poster

Model quantization is an effective way to compress deep neural networks and accelerate the inference time on edge devices. Existing quantization methods usually require original data for calibration during the compressing process, which may be inaccessible due to privacy issues. A common way is to g…

Cited by 1SourcePDFScholar
2025

Neural Variable-Order Fractional Differential Equation Networks

AAAI 2025technical

The use of neural differential equation models in machine learning applications has gained significant traction in recent years. In particular, fractional differential equations (FDEs) have emerged as a powerful tool for capturing complex dynamics in various domains. While existing models have prima…

Cited by 1SourcePDFScholar
2025

Rethinking Graph Neural Networks From A Geometric Perspective Of Node Features

ICLR 2025poster

Many works on graph neural networks (GNNs) focus on graph topologies and analyze graph-related operations to enhance performance on tasks such as node classification. In this paper, we propose to understand GNNs based on a feature-centric approach. Our main idea is to treat the features of nodes fro…

Cited by 0SourcePDFScholar
2025

Towards a General Time Series Anomaly Detector with Adaptive Bottlenecks and Dual Adversarial Decoders

ICLR 2025poster

Time series anomaly detection plays a vital role in a wide range of applications. Existing methods require training one specific model for each dataset, which exhibits limited generalization capability across different target datasets, hindering anomaly detection performance in various scenarios wit…

Cited by 5SourcePDFScholar
2025

Towards a General Time Series Forecasting Model with Unified Representation and Adaptive Transfer

ICML 2025poster

With the growing availability of multi-domain time series data, there is an increasing demand for general forecasting models pre-trained on multi-source datasets to support diverse downstream prediction scenarios. Existing time series foundation models primarily focus on scaling up pre-training data…

Cited by 0SourcePDFScholar
2025

Visual Autoregressive Modeling for Image Super-Resolution

ICML 2025poster

Image Super-Resolution (ISR) has seen significant progress with the introduction of remarkable generative models. However, challenges such as the trade-off issues between fidelity and realism, as well as computational complexity, have also posed limitations on their application. Building upon the tr…

2024

Chronic Poisoning: Backdoor Attack against Split Learning

AAAI 2024technical

Split learning is a computing resource-friendly distributed learning framework that protects client training data by splitting the model between the client and server. Previous work has proved that split learning faces a severe risk of privacy leakage, as a malicious server can recover the client's…

2024

Coupling Graph Neural Networks with Fractional Order Continuous Dynamics: A Robustness Study

AAAI 2024technical

In this work, we rigorously investigate the robustness of graph neural fractional-order differential equation (FDE) models. This framework extends beyond traditional graph neural (integer-order) ordinary differential equation (ODE) models by implementing the time-fractional Caputo derivative. Utiliz…

Cited by 6SourcePDFScholar
2024

DistilVPR: Cross-Modal Knowledge Distillation for Visual Place Recognition

AAAI 2024technical

The utilization of multi-modal sensor data in visual place recognition (VPR) has demonstrated enhanced performance compared to single-modal counterparts. Nonetheless, integrating additional sensors comes with elevated costs and may not be feasible for systems that demand lightweight operation, there…

2024

Distributed-Order Fractional Graph Operating Network

NeurIPS 2024spotlight

We introduce the Distributed-order fRActional Graph Operating Network (DRAGON), a novel continuous Graph Neural Network (GNN) framework that incorporates distributed-order fractional calculus. Unlike traditional continuous GNNs that utilize integer-order or single fractional-order differential equa…

2024

ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

IJCAI 2024poster

Offline reinforcement learning (RL) is a learning paradigm where an agent learns from a fixed dataset of experience. However, learning solely from a static dataset can limit the performance due to the lack of exploration. To overcome it, offline-to-online RL combines offline pre-training with online…

Cited by 6SourcePDFScholar
2024

Exploring Urban Semantics: A Multimodal Model for POI Semantic Annotation with Street View Images and Place Names

IJCAI 2024poster

Semantic annotation for points of interest (POIs) is the process of annotating a POI with a category label, which facilitates many services related to POIs, such as POI search and recommendation. Most of the existing solutions extract features related to POIs from abundant user-generated content dat…

2024

Learning Hierarchy-Enhanced POI Category Representations Using Disentangled Mobility Sequences

IJCAI 2024poster

Points of interest (POIs) carry a wealth of semantic information of varying locations in cities and thus have been widely used to enable various location-based services. To understand POI semantics, existing methods usually model contextual correlations of POI categories in users' check-in sequences…

2024

PosDiffNet: Positional Neural Diffusion for Point Cloud Registration in a Large Field of View with Perturbations

AAAI 2024technical

Point cloud registration is a crucial technique in 3D computer vision with a wide range of applications. However, this task can be challenging, particularly in large fields of view with dynamic objects, environmental noise, or other perturbations. To address this challenge, we propose a model called…

2024

Self-Promoted Clustering-based Contrastive Learning for Brain Networks Pretraining

IJCAI 2024poster

Rapid advancements in neuroimaging techniques, such as magnetic resonance imaging (MRI), have facilitated the acquisition of the structural and functional characteristics of the brain. Brain network analysis is one of the essential tools for exploring brain mechanisms from MRI, providing valuable in…

Cited by 0SourcePDFScholar
2024

Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback

ICLR 2024poster

Reinforcement Learning with Human Feedback (RLHF) has received significant attention for performing tasks without the need for costly manual reward design by aligning human preferences. It is crucial to consider diverse human feedback types and various learning methods in different environments. How…

2024

Unleashing the Potential of Fractional Calculus in Graph Neural Networks with FROND

ICLR 2024spotlight

We introduce the FRactional-Order graph Neural Dynamical network (FROND), a new continuous graph neural network (GNN) framework. Unlike traditional continuous GNNs that rely on integer-order differential equations, FROND employs the Caputo fractional derivative to leverage the non-local properties o…

2024

Urban Region Embedding via Multi-View Contrastive Prediction

AAAI 2024technical

Recently, learning urban region representations utilizing multi-modal data (information views) has become increasingly popular, for deep understanding of the distributions of various socioeconomic features in cities. However, previous methods usually blend multi-view information in a posteriors stag…

2024

XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution

ECCV 2024poster

"Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation inform…

2023

Adversarial Robustness in Graph Neural Networks: A Hamiltonian Approach

NeurIPS 2023spotlight

Graph neural networks (GNNs) are vulnerable to adversarial perturbations, including those that affect both node features and graph topology. This paper investigates GNNs derived from diverse neural flows, concentrating on their connection to various stability notions such as BIBO stability, Lyapunov…

2023

Graph Neural Convection-Diffusion with Heterophily

IJCAI 2023poster

Graph neural networks (GNNs) have shown promising results across various graph learning tasks, but they often assume homophily, which can result in poor performance on heterophilic graphs. The connected nodes are likely to be from different classes or have dissimilar features on heterophilic graphs.…

2023

HypLiLoc: Towards Effective LiDAR Pose Regression With Hyperbolic Fusion

CVPR 2023poster

LiDAR relocalization plays a crucial role in many fields, including robotics, autonomous driving, and computer vision. LiDAR-based retrieval from a database typically incurs high computation storage costs and can lead to globally inaccurate pose estimations if the database is too sparse. On the othe…

2023

Leveraging Label Non-Uniformity for Node Classification in Graph Neural Networks

ICML 2023poster

In node classification using graph neural networks (GNNs), a typical model generates logits for different class labels at each node. A softmax layer often outputs a label prediction based on the largest logit. We demonstrate that it is possible to infer hidden graph structural information from the d…

2023

Multi-Defendant Legal Judgment Prediction via Hierarchical Reasoning

EMNLP 2023long findings

Multiple defendants in a criminal fact description generally exhibit complex interactions, and cannot be well handled by existing Legal Judgment Prediction (LJP) methods which focus on predicting judgment results (e.g., law articles, charges, and terms of penalty) for single-defendant cases. To addr…

Cited by 0SourcecodeScholar
2023

Node Embedding from Neural Hamiltonian Orbits in Graph Neural Networks

ICML 2023poster

In the graph node embedding problem, embedding spaces can vary significantly for different data types, leading to the need for different GNN model types. In this paper, we model the embedding update of a node feature as a Hamiltonian orbit over time. Since the Hamiltonian orbits generalize the expon…

2023

Quality-Aware Pre-Trained Models for Blind Image Quality Assessment

CVPR 2023poster

Blind image quality assessment (BIQA) aims to automatically evaluate the perceived quality of a single image, whose performance has been improved by deep learning-based methods in recent years. However, the paucity of labeled data somewhat restrains deep learning-based BIQA methods from unleashing t…

Cited by 99SourcePDFScholar
2023

RPG-Palm: Realistic Pseudo-data Generation for Palmprint Recognition

ICCV 2023poster

Palmprint recently shows great potential in recognition applications as it is a privacy-friendly and stable biometric. However, the lack of large-scale public palmprint datasets limits further research and development of palmprint recognition. In this paper, we propose a novel realistic pseudo-palmp…

Cited by 12PDFScholar
2023

Towards an Integrated View of Semantic Annotation for POIs with Spatial and Textual Information

IJCAI 2023poster

Categories of Point of Interest (POI) facilitate location-based services from many aspects like location search and POI recommendation. However, POI categories are often incomplete and new POIs are being consistently generated, this rises the demand for semantic annotation for POIs, i.e., labeling t…

Cited by 9SourcePDFScholar
2022

A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation

IJCAI 2022poster

Reliable navigation systems have a wide range of applications in robotics and autonomous driving. Current approaches employ an open-loop process that converts sensor inputs directly into actions. However, these open-loop schemes are challenging to handle complex and dynamic real-world scenarios due…

2022

BézierPalm: A Free Lunch for Palmprint Recognition

ECCV 2022poster

"Palmprints are private and stable information for biometric recognition. In the deep learning era, the development of palmprint recognition is limited by the lack of sufficient training data. In this paper, by observing that palmar creases are the key information to deep-learning-based palmprint re…

Cited by 21SourcePDFScholar
2022

ContrastMask: Contrastive Learning To Segment Every Thing

CVPR 2022poster

Partially-supervised instance segmentation is a task which requests segmenting objects from novel categories via learning on limited base categories with annotated masks thus eliminating demands of heavy annotation burden. The key to addressing this task is to build an effective class-agnostic mask…

Cited by 52PDFcodeScholar
2022

Trajectory Tracking Control for Delta Parallel Manipulators: A Variable Gain ADRC Approach

RA-L 2022

In this letter, a novel variable gain active disturbance rejection control (VGADRC) approach is developed for solving the trajectory tracking problem of delta parallel manipulators with disturbances. The proposed VGADRC, including variable gain tracking differentiator (VGTD), variable gain extended

Cited by 12SourceScholar
2021

Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert Approach

IROS 2021poster

Despite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in…

Cited by 9SourceScholar
2021

Structured sparsification with joint optimization of group convolution and channel shuffle

UAI 2021poster

Recent advances in convolutional neural networks (CNNs) usually come with the expense of excessive computational overhead and memory footprint. Network compression aims to alleviate this issue by training compact models with comparable performance. However, existing compression techniques either ent…

2019

Optimizing the F-Measure for Threshold-Free Salient Object Detection

ICCV 2019poster

Current CNN-based solutions to salient object detection (SOD) mainly rely on the optimization of cross-entropy loss (CELoss). Then the quality of detected saliency maps is often evaluated in terms of F-measure. In this paper, we investigate an interesting issue: can we consistently use the F-measure…

Cited by 90PDFScholar
2019

Translate-to-Recognize Networks for RGB-D Scene Recognition

CVPR 2019poster

Cross-modal transfer is helpful to enhance modality-specific discriminative power for scene recognition. To this end, this paper presents a unified framework to integrate the tasks of cross-modal translation and modality-specific recognition, termed as Translate-to-Recognize Network TRecgNet. Specif…

Cited by 64PDFcodeScholar
2016

Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side Outputs

CVPR 2016poster

Object skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be…

Cited by 133PDFScholar