← Search

Yifan Zhao

31 accepted papers

2026

Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation Model

AAAI 2026technical

Cytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationsh

Cited by 0SourcePDFScholar
2026

Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

ICML 2026poster

The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypical settings. We term this failure as representation fragmentation, arising from a…

Cited by 0SourceScholar
2026

GSUC-VLM: Geometrically-Guided Spatial Understanding Chain of Vision Language Model for Autonomous Driving

ICRA 2026poster

Robust spatial understanding is crucial for Visual Question Answering (VQA) in autonomous driving that aims to enhance decision-making, reduce positional risks, and ensure road safety by providing answers based on the perception, prediction, and planning of driving scenarios. Despite remarkable succ…

Cited by 0Scholar
2026

Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events

CVPR 2026

Novel view synthesis from low dynamic range (LDR) blurry images, which are common in the wild, struggles to recover high dynamic range (HDR) and sharp 3D representations in extreme lighting conditions. Although existing methods employ event data to address this issue, they ignore the sensor-physics

Cited by 0SourcecodeScholar
2025

Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL

NeurIPS 2025poster

Few-Shot Class-Incremental Learning (FSCIL) challenges models to sequentially learn new classes from minimal examples without forgetting prior knowledge, a task complicated by the stability-plasticity dilemma and data scarcity. Current FSCIL methods often struggle with generalization due to their re…

Cited by 0SourceScholar
2025

Exploiting Motion Prior for Accurate Pose Estimation of Dashboard Cameras

RA-L 2025

Dashboard cameras (dashcams) record millions of driving videos daily, offering a valuable potential data source for various applications, including driving map production and updates. A necessary step for utilizing these dashcam data involves the estimation of camera poses. However, the low-quality

Cited by 1SourceScholar
2025

FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field

ICRA 2025

Visual SLAM has regained attention due to its ability to provide perceptual capabilities and simulation test data for Embodied AI. However, traditional SLAM methods struggle to meet the demands of high-quality scene reconstruction, and Gaussian SLAM systems, despite their rapid rendering and high-qu

Cited by 5SourceScholar
2025

FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation

ICCV 2025poster

Layout-to-image (L2I) generation has exhibited promising results in natural domains, but suffers from limited generative fidelity and weak alignment with user-provided layouts when applied to degraded scenes (i.e., low-light, underwater). We primarily attribute these limitations to the "contextual i…

Cited by 0SourcePDFScholar
2025

Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement

ICCV 2025poster

Class-Incremental Semantic Segmentation (CISS) requires continuous learning of newly introduced classes while retaining knowledge of past classes. By abstracting mainstream methods into two stages (visual feature extraction and prototype-feature matching), we identify a more fundamental challenge te…

Cited by 0SourcePDFScholar
2025

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

CVPR 2025poster

In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Model…

Cited by 0SourcePDFScholar
2025

Re-coding for Uncertainties: Edge-awareness Semantic Concordance for Resilient Event-RGB Segmentation

NeurIPS 2025poster

Semantic segmentation has achieved great success in ideal conditions. However, when facing extreme conditions (e.g., insufficient light, fierce camera motion), most existing methods suffer from significant information loss of RGB, severely damaging segmentation results. Several researches exploit th…

Cited by 0SourcecodeScholar
2025

When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network

ICML 2025spotlight

Anomaly detection is essential for the safety and reliability of autonomous driving systems. Current methods often focus on detection accuracy but neglect response time, which is critical in time-sensitive driving scenarios. In this paper, we introduce real-time anomaly detection for autonomous driv…

Cited by 0SourcePDFScholar
2024

Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RL

NeurIPS 2024poster

Accurate environment dynamics modeling is crucial for obtaining effective state representations in visual reinforcement learning (RL) applications. However, when facing multiple input modalities, existing dynamics modeling methods (e.g., DeepMDP) usually stumble in addressing the complex and volatil…

Cited by 0SourcePDFScholar
2024

SpikeNeRF: Learning Neural Radiance Fields from Continuous Spike Stream

CVPR 2024poster

Spike cameras leveraging spike-based integration sampling and high temporal resolution offer distinct advantages over standard cameras. However existing approaches reliant on spike cameras often assume optimal illumination a condition frequently unmet in real-world scenarios. To address this we intr…

2023

Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning

NeurIPS 2023poster

Integrating RGB frames with alternative modality inputs is gaining increasing traction in many vision-based reinforcement learning (RL) applications. Existing multi-modal vision-based RL methods usually follow a Global Value Estimation (GVE) pipeline, which uses a fused modality feature to obtain a…

2023

Learning With Fantasy: Semantic-Aware Virtual Contrastive Constraint for Few-Shot Class-Incremental Learning

CVPR 2023poster

Few-shot class-incremental learning (FSCIL) aims at learning to classify new classes continually from limited samples without forgetting the old classes. The mainstream framework tackling FSCIL is first to adopt the cross-entropy (CE) loss for training at the base session, then freeze the feature ex…

2023

Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement Learning

ICCV 2023accepted

Efficient motion and appearance modeling are critical for vision-based Reinforcement Learning (RL). However, existing methods struggle to reconcile motion and appearance information within the state representations learned from a single observation encoder. To address the problem, we present Synergi…

Cited by 1SourcePDFScholar
2023

Stabilizing Visual Reinforcement Learning via Asymmetric Interactive Cooperation

ICCV 2023poster

Vision-based reinforcement learning (RL) depends on discriminative representation encoders to abstract the observation states. Despite the great success of increasing CNN parameters for many supervised computer vision tasks, reinforcement learning with temporal-difference (TD) losses cannot benefit…

Cited by 4PDFScholar
2022

Spectrum Random Masking for Generalization in Image-based Reinforcement Learning

NeurIPS 2022accept

Generalization in image-based reinforcement learning (RL) aims to learn a robust policy that could be applied directly on unseen visual environments, which is a challenging task since agents usually tend to overfit to their training environment. To handle this problem, a natural approach is to incre…

Cited by 19SourcePDFScholar
2021

FACIAL: Synthesizing Dynamic Talking Face With Implicit Attribute Learning

ICCV 2021poster

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and eye blinks that are in-sync with the input audio signal. We…

Cited by 161PDFcodeScholar
2021

Heterogeneous Relational Complement for Vehicle Re-Identification

ICCV 2021poster

The crucial problem in vehicle re-identification is to find the same vehicle identity when reviewing this object from cross-view cameras, which sets a higher demand for learning viewpoint-invariant representations. In this paper, we propose to solve this problem from two aspects: constructing robust…

Cited by 69PDFcodeScholar
2021

Transformer-Based Dual Relation Graph for Multi-Label Image Recognition

ICCV 2021poster

The simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsistent appearances, and confused inter-class relationships. Recent research efforts mainly resort to the statistic label co…

Cited by 120PDFcodeScholar
2019

Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set Clustering

CVPR 2019poster

The reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically…

Cited by 13PDFScholar