← Search

Yanhao Ge

15 accepted papers

2026

LearnIR: Learnable Posterior Sampling for Real-World Image Restoration

ICLR 2026poster

Image restoration in real-world conditions is highly challenging due to heterogeneous degradations such as haze, noise, shadows, and blur. Existing diffusion-based methods remain limited: conditional generation struggles to balance fidelity and realism, inversion-based approaches accumulate errors,…

Cited by 0SourcecodeScholar
2026

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

CVPR 2026

Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet existing methods face inherent limitations: architecture-based approaches incur additional computational overhead and often generalize poorly to new tas

Cited by 0SourceScholar
2026

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

CVPR 2026

Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency texture details. We introduce VINS-120K, the first large-scale dataset for instruction-based UHR image editing, comprising 12

Cited by 0SourceScholar
2025

Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent

ICCV 2025poster

Despite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descripti…

Cited by 0SourcePDFScholar
2025

PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing

ICLR 2025poster

In the field of image editing, three core challenges persist: controllability, background preservation, and efficiency. Inversion-based methods rely on time-consuming optimization to preserve the features of the initial images, which results in low efficiency due to the requirement for extensive net…

2025

Toward Better Out-painting: Improving the Image Composition with Initialization Policy Model

ICCV 2025poster

With its extensive applications, Foreground Conditioned Out-painting (FCO) has attracted considerable attention in the research field. Through the utilization of text-driven FCO, users are enabled to generate diverse backgrounds for a given foreground by adjusting the text prompt, which considerably…

Cited by 0SourcePDFScholar
2025

UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset

NeurIPS 2025poster

Ultra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle t…

Cited by 0SourcecodeScholar
2024

Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

ECCV 2024poster

"Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and the results have not yet achieved satisfactory performance…

Cited by 27SourcePDFScholar
2024

HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation

ECCV 2024poster

"Recent advancements in text-to-image diffusion models have shown remarkable creative capabilities with textual prompts, but generating personalized instances based on specific subjects, known as subject-driven generation, remains challenging. To tackle this issue, we present a new hybrid framework…

Cited by 4SourcePDFScholar
2024

NeuMA: Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics

NeurIPS 2024poster

While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure neural-network-based simulators (black box), which may violate physical laws, or traditional physical simulators (white…

2022

Learning To Restore 3D Face From In-the-Wild Degraded Images

CVPR 2022poster

In-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsuperv…

Cited by 3PDFScholar
2022

Physically-Guided Disentangled Implicit Rendering for 3D Face Modeling

CVPR 2022poster

This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for high-fidelity 3D face modeling. The motivation comes from two observations: widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering metho…

Cited by 8PDFScholar
2021

Learning To Aggregate and Personalize 3D Face From In-the-Wild Photo Collection

CVPR 2021poster

Non-prior face modeling aims to reconstruct 3D face only from images without shape assumptions. While plausible facial details are predicted, the models tend to over-depend on local color appearance and suffer from ambiguous noise. To address such problem, this paper presents a novel Learning to Agg…

Cited by 34PDFScholar
2020

Adversarial Semantic Data Augmentation for Human Pose Estimation

ECCV 2020poster

Human pose estimation is the task of localizing body keypoints from still images. The state-of-the-art methods suffer from insufficient examples of challenging cases such as symmetric appearance, heavy occlusion and nearby person. To enlarge the amounts of challenging cases, previous methods augment…