← Search

Xiaoyan Sun

41 accepted papers

2026

Content-Adaptive Hierarchical Hyperprior for Neural Video Coding

CVPR 2026

While neural video codecs (NVCs) have recently demonstrated superior performance over traditional codecs through end-to-end learning, existing approaches primarily focus on architectural enhancements and coding module design, with limited exploration into optimizing hierarchical structures--specific

Cited by 0SourceScholar
2026

EvReflection: Event-Driven Micro-Dynamics for Reflection Removal

ICML 2026poster

Reflection removal is a highly challenging problem. Though remarkable progress has been made, current methods primarily exploit static image priors from a single frame. Due to the inherent ambiguity between layers, existing methods still suffer from severe residual artifacts. In this paper, we propo…

Cited by 0SourceScholar
2026

OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

ICML 2026poster

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic …

Cited by 0SourceScholar
2026

ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation

ICLR 2026poster

Personalized video generation faces a fundamental trade-off between identity consistency and action realism: overly rigid identity preservation often leads to unnatural motion, while emphasis on action dynamics can compromise subject fidelity. This tension stems from three interrelated challenges: i…

Cited by 0SourceScholar
2026

SEMAMIL: SEMANTIC-AWARE MULTIPLE INSTANCE LEARNING WITH RETRIEVAL-GUIDED STATE SPACE MODELING FOR WHOLE SLIDE IMAGES

ICASSP 2026poster

Multiple instance learning (MIL) has become the leading approach for extracting discriminative features from whole slide images (WSIs) in computational pathology. Attention-based MIL methods can identify key patches but tend to overlook contextual relationships. Transformer models are able to model…

Cited by 0SourcePDFScholar
2026

SSCM: A SPATIAL-SEMANTIC CONSISTENT MODEL FOR MULTI-CONTRAST MRI SUPER-RESOLUTION

ICASSP 2026poster

Multi-contrast Magnetic Resonance Imaging super-resolution (MC-MRI SR) aims to enhance low-resolution (LR) contrasts leveraging high-resolution (HR) references, shortening acquisition time and improving imaging efficiency while preserving anatomical details. The main challenge lies in maintaining sp…

Cited by 0SourcePDFScholar
2026

Seeing the Unseen: Zooming in the Dark with Event Cameras

AAAI 2026technical

This paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these chall

Cited by 0SourcePDFScholar
2026

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

ICML 2026poster

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks …

Cited by 0SourceScholar
2025

D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement

ICLR 2025spotlight

We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation…

2025

DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering

ICCV 2025poster

Dynamic scene reconstruction is a long-term challenge in 3D vision. Existing plane-based methods in dynamic Gaussian splatting suffer from an unsuitable low-rank assumption, causing feature overlap and poor rendering quality. Although 4D hash encoding provides an explicit representation without low-…

2025

Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach

AAAI 2025technical

Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events…

2025

Efficient Spiking Point Mamba for Point Cloud Analysis

ICCV 2025poster

Bio-inspired Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing 3D SNNs have struggled with long-range dependencies until the recent emergence of Mamba, which offers superior computational efficiency and sequence modeling capabili…

2025

Event-Enhanced Blurry Video Super-Resolution

AAAI 2025technical

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insu…

2025

Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model

CVPR 2025poster

Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of availabl…

Cited by 0SourcePDFScholar
2025

RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis

EMNLP 2025

Retrieval-Augmented Generation (RAG) enriches the input to LLMs by retrieving information from the relevant knowledge database, enabling them to produce responses that are more accurate and contextually appropriate. It is worth noting that the knowledge database, being sourced from publicly availabl

Cited by 0SourcePDFScholar
2025

Spiking Point Transformer for Point Cloud Classification

AAAI 2025technical

Spiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud rem…

2024

EvTexture: Event-driven Texture Enhancement for Video Super-Resolution

ICML 2024poster

Event-based vision has drawn increasing attention due to its unique characteristics, such as high temporal resolution and high dynamic range. It has been used in video super-resolution (VSR) recently to enhance the flow estimation and temporal alignment. Rather than for motion learning, we propose i…

2024

Event-Adapted Video Super-Resolution

ECCV 2024poster

"Introducing event cameras into video super-resolution (VSR) shows great promise. In practice, however, integrating event data as a new modality necessitates a laborious model architecture design. This not only consumes substantial time and effort but also disregards valuable insights from successfu…

Cited by 6SourcePDFScholar
2024

Event-assisted Low-Light Video Object Segmentation

CVPR 2024poster

In the realm of video object segmentation (VOS) the challenge of operating under low-light conditions persists resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras characterized by their high dynamic ran…

2024

Event-based Head Pose Estimation: Benchmark and Method

ECCV 2024poster

"Head pose estimation (HPE) is crucial for various applications, including human-computer interaction, augmented reality, and driver monitoring. However, traditional RGB-based methods struggle in challenging conditions like sudden movement and extreme lighting. Event cameras, as a neuromorphic senso…

2024

Exploiting Dual-Correlation for Multi-frame Time-of-Flight Denoising

ECCV 2024poster

"Recent advancements in Time-of-Flight (ToF) depth denoising have achieved impressive results in removing Multi-Path Interference (MPI) and shot noise. However, existing methods only utilize a single frame of ToF data, neglecting the correlation between frames. In this paper, we propose the first le…

2024

Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced Learning

ICASSP 2024accepted

Low-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. More…

Cited by 0SourceScholar
2024

Image Captioning with Multi-Context Synthetic Data

AAAI 2024technical

Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This potential can be harnessed to create synthetic image-text p…

Cited by 13SourcePDFScholar
2024

MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation

CVPR 2024highlight

We present MicroCinema a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage proces…

Cited by 15SourcePDFScholar
2024

Panacea: Panoramic and Controllable Video Generation for Autonomous Driving

CVPR 2024poster

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper we propose Panacea an innovative approach to generate panoramic and controllable videos in driving scenarios capable of yielding an unlimited numbers of diverse annotated samples pivotal for auto…

Cited by 50SourcePDFScholar
2024

Scene Adaptive Sparse Transformer for Event-based Object Detection

CVPR 2024poster

While recent Transformer-based approaches have shown impressive performances on event-based object detection tasks their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. H…

2024

TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities

AAAI 2024technical

Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes…

Cited by 5SourcePDFScholar
2024

Visual Perception by Large Language Model’s Weights

NeurIPS 2024poster

Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstra…

2023

Better and Faster: Adaptive Event Conversion for Event-Based Object Detection

AAAI 2023technical

Event cameras are a kind of bio-inspired imaging sensor, which asynchronously collect sparse event streams with many advantages. In this paper, we focus on building better and faster event-based object detectors. To this end, we first propose a computationally efficient event representation Hyper Hi…

Cited by 18SourcePDFScholar
2023

GET: Group Event Transformer for Event-Based Vision

ICCV 2023poster

Event cameras are a type of novel neuromorphic sen-sor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity.…

Cited by 107PDFcodeScholar
2023

Paint by Example: Exemplar-Based Image Editing With Diffusion Models

CVPR 2023poster

Language-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive ap…

2022

Exploiting Rigidity Constraints for LiDAR Scene Flow Estimation

CVPR 2022poster

Previous LiDAR scene flow estimation methods, especially recurrent neural networks, usually suffer from structure distortion in challenging cases, such as sparse reflection and motion occlusions. In this paper, we propose a novel optimization method based on a recurrent neural network to predict LiD…

Cited by 40PDFScholar
2021

Dual Progressive Prototype Network for Generalized Zero-Shot Learning

NeurIPS 2021poster

Generalized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain tran…

Cited by 57SourcePDFScholar
2021

Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning

AAAI 2021technical

Generalized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. How…

Cited by 18SourcePDFScholar
2021

Training Spiking Neural Networks with Accumulated Spiking Flow

AAAI 2021technical

The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demon…

2018

MiCT: Mixed 3D/2D Convolutional Tube for Human Action Recognition

CVPR 2018poster

Human actions in videos are three-dimensional (3D) signals. Recent attempts use 3D convolutional neural networks (CNNs) to explore spatio-temporal information for human action recognition. Though promising, 3D CNNs have not achieved high performanceon on this task with respect to their well-establis…

Cited by 295SourcePDFScholar