← Search

Zeyu Xiao

19 accepted papers

2026

FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation

AAAI 2026technical

Layout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, w

Cited by 0SourcePDFScholar
2026

KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals

AAAI 2026technical

Full-body motion tracking plays an essential role in AR/VR applications, bridging physical and virtual interactions. However, it is challenging to reconstruct realistic and diverse full-body poses based on sparse signals obtained by head-mounted displays, which are the main devices in AR/VR scenario

Cited by 0SourcePDFScholar
2026

Robust Vision-Language Models via Manifold-Adversarial Adapters

ICML 2026poster

Vision-language models (VLMs) have progressed rapidly with large-scale high-quality data and adaptation strategies, yet remain brittle under real-world corruptions, where both visual recognition and language-grounded reasoning degrade. Beyond cascaded image restoration, a natural alternative is para…

Cited by 0SourceScholar
2026

Seeing the Unseen: Zooming in the Dark with Event Cameras

AAAI 2026technical

This paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these chall

Cited by 0SourcePDFScholar
2026

UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation

AAAI 2026technical

Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately,

Cited by 0SourcePDFScholar
2025

Asymmetric Dual-Lens Video Deblurring

NeurIPS 2025poster

Modern smartphones often feature asymmetric dual-lens systems, capturing wide-angle and ultra-wide views with complementary perspectives and details. Motion and shake can blur the wide lens, while the ultra-wide lens, despite lower resolution, retains sharper details. This natural complementarity of…

Cited by 0SourceScholar
2025

Event-Enhanced Blurry Video Super-Resolution

AAAI 2025technical

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insu…

2024

Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models

ICML 2024poster

Fine-tuning pre-trained vision-language models (VLMs), e.g., CLIP, for the open-world generalization has gained increasing popularity due to its practical value. However, performance advancements are limited when relying solely on intricate algorithmic designs for a single model, even one exhibiting…

2024

Event-Adapted Video Super-Resolution

ECCV 2024poster

"Introducing event cameras into video super-resolution (VSR) shows great promise. In practice, however, integrating event data as a new modality necessitates a laborious model architecture design. This not only consumes substantial time and effort but also disregards valuable insights from successfu…

Cited by 6SourcePDFScholar
2024

Learning Large-Factor EM Image Super-Resolution with Generative Priors

CVPR 2024poster

As the mainstream technique for capturing images of biological specimens at nanometer resolution electron microscopy (EM) is extremely time-consuming for scanning wide field-of-view (FOV) specimens. In this paper we investigate a challenging task of large-factor EM image super-resolution (EMSR) whic…

2023

CutMIB: Boosting Light Field Super-Resolution via Multi-View Image Blending

CVPR 2023poster

Data augmentation (DA) is an efficient strategy for improving the performance of deep neural networks. Recent DA strategies have demonstrated utility in single image super-resolution (SR). Little research has, however, focused on the DA strategy for light field SR, in which multi-view information ut…

2022

Efficient Model-Driven Network for Shadow Removal

AAAI 2022technical

Deep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant prope…

2022

Frequency and Spatial Dual Guidance for Image Dehazing

ECCV 2022poster

"In this paper, we propose a novel image dehazing framework with frequency and spatial dual guidance. In contrast to most existing deep learning-based image dehazing methods that primarily exploit spatial information and neglect the distinguished frequency information, we introduce a new perspective…

2021

Unfolding Taylor's Approximations for Image Restoration

NeurIPS 2021poster

Deep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping ne…

Cited by 26SourcePDFScholar