← Search

Ao Luo

28 accepted papers

2026

DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesis

CVPR 2026

Image alignment is a fundamental task in computer vision with broad applications. Existing methods predominantly employ optical flow-based image warping. However, this technique is susceptible to common challenges such as occlusions and illumination variations, leading to degraded alignment visual q

Cited by 0SourcecodeScholar
2026

GeneVAR: Causal MeanFlow for Autoregressive Gene-to-WSI Tile Synthesis

CVPR 2026

Understanding how transcriptomic programs shape tissue morphology remains a central challenge in computational pathology. Gene-to-WSI tile synthesis offers a principled generative framework to translate molecular profiles into histological images. However, most existing methods compress RNA-Seq into

Cited by 0SourceScholar
2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

ICLR 2026poster

Despite substantial progress in video understanding, most existing datasets are limited to Earth’s gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-rob…

Cited by 0SourcecodeScholar
2026

I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle

AAAI 2026technical

Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives,

Cited by 0SourcePDFScholar
2026

Optical Flow Matching: Reframing Optical Flow as Continuous Transport Dynamics

CVPR 2026

Modern optical flow estimation, though empowered by recent deep neural architectures, remains rooted in the discrete correspondence paradigm inherited from classical vision. Most networks infer frame-to-frame displacements, capturing where pixels move but not how motion evolves continuously through

Cited by 0SourcecodeScholar
2025

Forensics Adapter: Adapting CLIP for Generalizable Face Forgery Detection

CVPR 2025poster

We describe the Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is non-trivial as forgery-related knowledge is entangled with a wide range of unrelate…

Cited by 5SourcePDFScholar
2025

MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification

CVPR 2025poster

Whole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the str…

Cited by 0SourcePDFScholar
2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

NeurIPS 2024poster

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and ac…

2024

Efficient Meshflow and Optical Flow Estimation from Event Cameras

CVPR 2024poster

In this paper we explore the problem of event-based meshflow estimation a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start we generate a large-scale High-Resolution Event Meshflow (HREM) dataset which showcases its superiority by encompassing th…

2024

FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models

CVPR 2024highlight

Optical flow estimation a process of predicting pixel-wise displacement between consecutive frames has commonly been approached as a regression task in the age of deep learning. Despite notable advancements this de facto paradigm unfortunately falls short in generalization performance when trained o…

2024

FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection

ECCV 2024poster

"Detecting objects seamlessly blended into their surroundings represents a complex task for both human cognitive capabilities and advanced artificial intelligence algorithms. Currently, the majority of methodologies for detecting camouflaged objects mainly focus on utilizing discriminative models wi…

2024

LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models

ECCV 2024poster

"In this paper, we propose a diffusion-based unsupervised framework that incorporates physically explainable Retinex theory with diffusion models for low-light image enhancement, named LightenDiffusion. Specifically, we present a content-transfer decomposition network that performs Retinex decomposi…

2024

RecDiffusion: Rectangling for Image Stitching with Diffusion Models

CVPR 2024poster

Image stitching from different captures often results in non-rectangular boundaries which is often considered unappealing. To solve non-rectangular boundaries current solutions involve cropping which discards image content inpainting which can introduce unrelated content or warping which can distort…

2024

SCP: Spherical-Coordinate-Based Learned Point Cloud Compression

AAAI 2024technical

In recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, LiDAR point cloud, is generated by spinning LiDAR on vehicles. This process results in numerous circular shapes and azimuthal angle invariance features within the point clouds. Howev…

2023

Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training

ACL 2023long

In this paper, we introduce Cross-View Language Modeling, a simple and effective pre-training framework that unifies cross-lingual and cross-modal pre-training with shared architectures and objectives. Our approach is motivated by a key observation that cross-lingual and cross-modal pre-training sha…

2023

Explicit Motion Disentangling for Efficient Optical Flow Estimation

ICCV 2023poster

In this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer t…

Cited by 18PDFcodeScholar
2023

GAFlow: Incorporating Gaussian Attention into Optical Flow

ICCV 2023poster

Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and s…

Cited by 32PDFcodeScholar
2023

Learning Optical Flow from Event Camera with Rendered Dataset

ICCV 2023poster

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted…

Cited by 19PDFcodeScholar
2022

Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras

IROS 2022poster

Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of…

Cited by 0SourceScholar
2022

Learning Optical Flow with Adaptive Graph Reasoning

AAAI 2022technical

Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explici…

2022

RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos

ECCV 2022poster

"Obtaining the ground truth labels from a video is challenging since the manual annotation of pixel-wise flow labels is prohibitively expensive and laborious. Besides, existing approaches try to adapt the trained model on synthetic datasets to authentic videos, which inevitably suffers from domain d…

2021

Probabilistic Model Distillation for Semantic Correspondence

CVPR 2021poster

Semantic correspondence is a fundamental problem in computer vision, which aims at establishing dense correspondences across images depicting different instances under the same category. This task is challenging due to large intra-class variations and a severe lack of ground truth. A popular solutio…

Cited by 26PDFcodeScholar
2021

TemporalFusion: Temporal Motion Reasoning with Multi-Frame Fusion for 6D Object Pose Estimation

IROS 2021poster

6D object pose estimation is an essential task in vision-based robotic grasping and manipulation. Prior works extract spatial features by fusing the RGB image and depth without considering the temporal motion information, limiting their performance in heavy occlusion robotic grasping scenarios. In t…

Cited by 4SourcecodeScholar
2021

Uncertainty-Guided Transformer Reasoning for Camouflaged Object Detection

ICCV 2021poster

Spotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones wi…

Cited by 296PDFcodeScholar
2021

WebSRC: A Dataset for Web-Based Structural Reading Comprehension

EMNLP 2021main

Web search is an essential way for humans to obtain information, but it’s still a great challenge for machines to understand the contents of web pages. In this paper, we introduce the task of web-based structural reading comprehension. Given a web page and a question about it, the task is to find an…

Cited by 85SourcePDFScholar
2020

Cascade Graph Neural Networks for RGB-D Salient Object Detection

ECCV 2020poster

In this paper, we study the problem of salient object detection for RGB-D images by using both color and depth information. A major technical challenge for detecting salient objects in RGB-D images is to fully leverage the two complementary data sources. The existing works either simply distill prio…

2019

End-to-End Driving Model for Steering Control of Autonomous Vehicles with Future Spatiotemporal Features

IROS 2019poster

End-to-end deep learning has gained considerable interests in autonomous driving vehicles in both academic and industrial fields, especially in decision making process. One critical issue in decision making process of autonomous driving vehicles is steering control. Researchers has already trained d…

Cited by 43SourceScholar