← Search

Yingnan Zhao

6 accepted papers

2026

DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor Attacks

AAAI 2026technical

Contrastive learning (CL) is a popular learning paradigm that excels in extracting meaningful representations from unlabeled data. Recent studies have shown that CL is highly vulnerable to backdoor attacks. Current defenses against backdoor attacks in CL are primarily reactive and post-training. Tha

Cited by 0SourcePDFScholar
2026

Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning

AAAI 2026technical

Humanoid robots are promising to learn a diverse set of human-like locomotion behaviors, including standing up, walking, running, and jumping. However, existing methods predominantly require training independent policies for each skill, yielding behavior-specific controllers that exhibit limited gen

Cited by 0SourcePDFScholar
2026

X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation

RSS 2026poster

While recent advances have demonstrated strong performance in individual humanoid skills such as upright locomotion, fall recovery and whole-body coordination, learning a single policy that masters all these skills remains challenging due to the diverse dynamics and conflicting control objectives in…

Cited by 0SourceScholar
2025

Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement Learning

NeurIPS 2025poster

The remarkable empirical performance of distributional reinforcement learning~(RL) has garnered increasing attention to understanding its theoretical advantages over classical RL. By decomposing the categorical distributional loss commonly employed in distributional RL, we find that the potential su…

Cited by 0SourceScholar
2024

Distributional Reinforcement Learning with Regularized Wasserstein Loss

NeurIPS 2024poster

The empirical success of distributional reinforcement learning (RL) highly relies on the choice of distribution divergence equipped with an appropriate distribution representation. In this paper, we propose \textit{Sinkhorn distributional RL (SinkhornDRL)}, which leverages Sinkhorn divergence—a regu…

2021

Damped Anderson Mixing for Deep Reinforcement Learning: Acceleration, Convergence, and Stabilization

NeurIPS 2021poster

Anderson mixing has been heuristically applied to reinforcement learning (RL) algorithms for accelerating convergence and improving the sampling efficiency of deep RL. Despite its heuristic improvement of convergence, a rigorous mathematical justification for the benefits of Anderson mixing in RL ha…

Cited by 19SourcePDFScholar