← Search

Yaozhong Gan

10 accepted papers

2026

A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World

CVPR 2026

Existing methods for deepfake detection aim to develop generalizable detectors. Although "generalizable" could be the ultimate target once and for all, with limited training forgeries and domains, it appears idealistic to expect generalization that covers entirely unseen variations, especially given

Cited by 1SourceScholar
2026

MARPO: A Reflective Policy Optimization for Multi-Agent Reinforcement Learning

AAAI 2026technical

We propose Multi-Agent Reflective Policy Optimization MARPO to alleviate the issue of sample inefficiency in multi-agent reinforcement learning. MARPO consists of two key components: a reflection mechanism that leverages subsequent trajectories to enhance sample efficiency, and an asymmetric clippin

Cited by 0SourcePDFScholar
2026

Reexamining the Exploration–Exploitation Dilemma from an Entropy-Driven Perspective

IJCAI 2026

Achieving an optimal balance between exploration and exploitation remains a fundamental challenge in reinforcement learning. This work revisits the exploration-exploitation dilemma through the lens of entropy, offering a novel perspective on this enduring problem. It establishes a theoretical connec

Cited by 0Scholar
2025

Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment

ICCV 2025poster

While fine-tuning diffusion models with reinforcement learning (RL) has demonstrated effectiveness in directly optimizing downstream objectives, existing RL frameworks are prone to overfitting the rewards, leading to outputs that deviate from the true data distribution and exhibit reduced diversity.…

Cited by 0SourcePDFScholar
2024

PAE: Reinforcement Learning from External Knowledge for Efficient Exploration

ICLR 2024poster

Human intelligence is adept at absorbing valuable insights from external knowledge. This capability is equally crucial for artificial intelligence. In contrast, classical reinforcement learning agents lack such capabilities and often resort to extensive trial and error to explore the environment.…

Cited by 2SourcePDFScholar