← Search

Congying Han

17 accepted papers

2026

On the Tension Between Optimality and Adversarial Robustness in Policy Optimization

ICLR 2026poster

Achieving optimality and adversarial robustness in deep reinforcement learning has long been regarded as conflicting goals. Nonetheless, recent theoretical insights presented in CAR suggest a potential alignment, raising the important question of how to realize this in practice. This paper first ide…

Cited by 0SourceScholar
2026

Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment

CVPR 2026

Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing methods learn this bridge by mimicking the score-matching formulation of standard diffusion models. In this work, we find that this way leads to an an

Cited by 0SourcecodeScholar
2025

A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form Games

ICML 2025poster

Solving the Nash equilibrium in normal-form games with large-scale strategy spaces presents significant challenges. Open-ended learning frameworks, such as PSRO and its variants, have emerged as effective solutions. However, these methods often lack an efficient metric for evaluating strategy improv…

Cited by 0SourcePDFScholar
2025

DreamHA: Towards High-Quality Human Animation with Image-to-Video Diffusion Models

ICASSP 2025accepted

Recent diffusion models have made significant advancements in generating lifelike videos from driving signals, including a reference character and a skeleton sequence. Nevertheless, these models often struggle with maintaining fidelity, as the generated results frequently deviate in character featur…

Cited by 0SourceScholar
2025

Feature out! Let Raw Image as Your Condition for Blind Face Restoration

ICML 2025poster

Blind face restoration (BFR), which involves converting low-quality (LQ) images into high-quality (HQ) images, remains challenging due to complex and unknown degradations. While previous diffusion-based methods utilize feature extractors from LQ images as guidance, using raw LQ images directly…

Cited by 0SourcePDFScholar
2025

MIRROR: Make Your Object-Level Multi-View Generation More Consistent with Training-Free Rectification

ICML 2025poster

Multi-view Diffusion has greatly advanced the development of 3D content creation by generating multiple images from distinct views, achieving remarkable photorealistic results. However, existing works are still vulnerable to inconsistent 3D geometric structures (commonly known as Janus Problem) and…

Cited by 0SourcePDFScholar
2025

Purity Law for Neural Routing Problem Solvers with Enhanced Generalizability

NeurIPS 2025poster

Achieving generalization in neural approaches across different scales and distributions remains a significant challenge for routing problems. A key obstacle is that neural networks often fail to learn robust principles for identifying universal patterns and deriving optimal solutions from diverse in…

Cited by 0SourceScholar
2025

StyO: Stylize Your Face in Only One-Shot

AAAI 2025technical

This paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here, we present a novel StyO model, i.e., Stylize the face in only One-shot, to solve the above problem. In particular, StyO…

Cited by 9SourcePDFScholar
2024

BlazeBVD: Make Scale-Time Equalization Great Again for Blind Video Deflickering

ECCV 2024poster

"Developing blind video deflickering (BVD) algorithms to enhance video temporal consistency, is gaining importance amid the flourish of image processing and video generation. However, the intricate nature of video data complicates the training of deep learning methods, leading to high resource consu…

Cited by 0SourcePDFScholar
2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

CVPR 2024poster

Recent works in implicit representations such as Neural Radiance Fields (NeRF) have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters since the lack of explicit geometric constraints pose…

2024

Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error

ICML 2024oral

Establishing robust policies is essential to counter attacks or disturbances affecting deep reinforcement learning (DRL) agents. Recent studies explore state-adversarial robustness and suggest the potential lack of an optimal robust policy (ORP), posing challenges in setting strict robustness constr…

2023

Towards Consistent Video Editing with Text-to-Image Diffusion Models

NeurIPS 2023poster

Existing works have advanced Text-to-Image (TTI) diffusion models for video editing in a one-shot learning manner. Despite their low requirements of data and computation, these methods might produce results of unsatisfied consistency with text prompt as well as temporal sequence, limiting their appl…

Cited by 32SourcePDFScholar
2023

Transforming Radiance Field With Lipschitz Network for Photorealistic 3D Scene Stylization

CVPR 2023highlight

Recent advances in 3D scene representation and novel view synthesis have witnessed the rise of Neural Radiance Fields (NeRFs). Nevertheless, it is not trivial to exploit NeRF for the photorealistic 3D scene stylization task, which aims to generate visually consistent and photorealistic stylized scen…

2022

Generalized One-shot Domain Adaptation of Generative Adversarial Networks

NeurIPS 2022accept

The adaptation of a Generative Adversarial Network (GAN) aims to transfer a pre-trained GAN to a target domain with limited training data. In this paper, we focus on the one-shot case, which is more challenging and rarely explored in previous works. We consider that the adaptation from a source doma…

2022

PetsGAN: Rethinking Priors for Single Image Generation

AAAI 2022technical

Single image generation (SIG), described as generating diverse samples that have the same visual content as the given natural image, is first introduced by SinGAN, which builds a pyramid of GANs to progressively learn the internal patch distribution of the single image. It shows excellent performanc…

2022

Shrinking Temporal Attention in Transformers for Video Action Recognition

AAAI 2022technical

Spatiotemporal modeling in an unified architecture is key for video action recognition. This paper proposes a Shrinking Temporal Attention Transformer (STAT), which efficiently builts spatiotemporal attention maps considering the attenuation of spatial attention in short and long temporal sequences.…

Cited by 13SourcePDFScholar