← Search

Yi Xiao

21 accepted papers

2026

A Hybrid Space Model for Misaligned Multi-modality Image Fusion

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing

Cited by 1SourcePDFScholar
2026

Content-aware Information Compression and Selection for Whole Slide Image Analysis

AAAI 2026technical

Recent advances in multi-instance learning (MIL) have demonstrated impressive performance in whole slide image (WSI) analysis. However, current methods search for cues and draw conclusions from all instances or regions, resulting in excessive redundant computation and suboptimal representation quali

Cited by 0SourcePDFScholar
2026

Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining

AAAI 2026technical

Rain significantly degrades the performance of computer vision systems, particularly in applications like autonomous driving and video surveillance. While existing deraining methods have made considerable progress, they often struggle with fidelity of semantic and spatial details. To address these l

Cited by 0SourcePDFScholar
2025

Designing Cyclic Peptides via Harmonic SDE with Atom-Bond Modeling

ICML 2025poster

Cyclic peptides offer inherent advantages in pharmaceuticals. For example, cyclic peptides are more resistant to enzymatic hydrolysis compared to linear peptides and usually exhibit excellent stability and affinity. Although deep generative models have achieved great success in linear peptide design…

Cited by 0SourcePDFScholar
2025

GMMamba: Group Masking Mamba for Whole Slide Image Classification

ICCV 2025poster

Recent advances in selective state space models (Mamba) have shown great promise in whole slide image (WSI) classification. Despite this, WSIs contain explicit local redundancy (similar patches) and irrelevant regions (uninformative instances), posing significant challenges for Mamba-based multi-ins…

Cited by 0SourcePDFScholar
2025

Integrating Protein Dynamics into Structure-Based Drug Design via Full-Atom Stochastic Flows

ICLR 2025poster

The dynamic nature of proteins, influenced by ligand interactions, is essential for comprehending protein function and progressing drug discovery. Traditional structure-based drug design (SBDD) approaches typically target binding sites with rigid structures, limiting their practical application in d…

Cited by 0SourcePDFScholar
2025

LoRA-EnVar: Parameter-Efficient Hybrid Ensemble Variational Assimilation for Weather Forecasting

NeurIPS 2025poster

Accurate estimation of background error (i.e., forecast error) distribution is critical for effective data assimilation (DA) in numerical weather prediction (NWP). In state-of-the-art operational DA systems, it is common to account for the temporal evolution of background errors by employing hybrid…

Cited by 0SourceScholar
2025

M3amba: Memory Mamba is All You Need for Whole Slide Image Classification

CVPR 2025poster

Multi-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solu…

Cited by 0SourcePDFScholar
2025

OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual Learning

AAAI 2025technical

Bag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labe…

Cited by 0SourcePDFScholar
2025

Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural Networks

NeurIPS 2025poster

Spiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain un…

Cited by 0SourcecodeScholar
2025

VAE-Var: Variational Autoencoder-Enhanced Variational Methods for Data Assimilation in Meteorology

ICLR 2025poster

Data assimilation (DA) is an essential statistical technique for generating accurate estimates of a physical system's states by combining prior model predictions with observational data, especially in the realm of weather forecasting. Effectively modeling the prior distribution while adapting to div…

Cited by 1SourcePDFScholar
2024

Towards a Self-contained Data-driven Global Weather Forecasting Framework

ICML 2024poster

Data-driven weather forecasting models are advancing rapidly, yet they rely on initial states (i.e., analysis states) typically produced by traditional data assimilation algorithms. Four-dimensional variational assimilation (4DVar) is one of the most widely adopted data assimilation algorithms in nu…

Cited by 8SourcePDFScholar
2024

VeriCompress: A Tool to Streamline the Synthesis of Verified Robust Compressed Neural Networks from Scratch

AAAI 2024technical

AI's widespread integration has led to neural networks (NN) deployment on edge and similar limited-resource platforms for safety-critical scenarios. Yet, NN's fragility raises concerns about reliable inference. Moreover, constrained platforms demand compact networks. This study introduces VeriCompre…

2023

Scaling Vision-Based End-to-End Autonomous Driving with Multi-View Attention Learning

IROS 2023poster

On end-to-end driving, human driving demonstrations are used to train perception-based driving models by imitation learning. This process is supervised on vehicle signals (e.g., steering angle, acceleration) but does not require extra costly supervision (human labeling of sensor data). As a represen…

Cited by 5SourceScholar
2023

Semantic-Aware Gated Fusion Network For Interactive Colorization

ICASSP 2023accepted

Deep neural networks boost many successful colorization methods, including automatic, interactive, and exemplar-based methods. Among them, interactive methods with global and/or local inputs are probably the most flexible to accurately add colors to a gray image. However, due to the sparseness of in…

Cited by 0SourceScholar
2022

Histogram-Guided Semantic-Aware Colorization

ICASSP 2022accepted

User-guided colorization can predict the colors of a grayscale image according to user inputs, including exemplar images, local inputs and global inputs. Global inputs-based methods are probably the easiest ones to use, but are hard to distribute the input colors into correct regions, due to the lac…

Cited by 0SourceScholar
2020

Action-based Representation Learning for Autonomous Driving

CoRL 2020

Human drivers produce a vast amount of data which could, in principle, be used to improve autonomous driving systems. Unfortunately, seemingly straightforward approaches for creating end-to-end driving models that map sensor data directly into driving actions are problematic in terms of interpretabi

2020

Sketchppnet: A Joint Pixel and Point Convolutional Neural Network For Low Resolution Sketch Image Recognition

ICASSP 2020accepted

Sketch recognition using deep neural networks have become a recent trend. However, traditional pixel (image) based convolutional neural networks show poor recognizing performance on low resolution (LR) sketch image due to the loss of image details. To solve this problem, we propose a joint pixel and…

Cited by 0SourceScholar
2019

Interactive Deep Colorization Using Simultaneous Global and Local Inputs

ICASSP 2019accepted

Colorization methods using deep neural networks have become a recent trend. However, most of them do not allow user inputs, or only allow limited user inputs (only global inputs or only local inputs), to control the output colorful images. The possible reason is that it's difficult to differentiate…

Cited by 0SourceScholar