← Search

Yuxuan Han

17 accepted papers

2026

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty

ICLR 2026poster

Deep reinforcement learning (RL) has achieved remarkable success, yet its deployment in real-world scenarios is often limited by vulnerability to environmental uncertainties. Distributionally robust RL (DR-RL) algorithms have been proposed to resolve this challenge, but existing approaches are large…

Cited by 0SourcecodeScholar
2026

Doubly Robust Distributionally Robust Offline Contextual Pricing

ICML 2026poster

Offline contextual pricing often relies on logged observational data, but faces challenges from distributional shifts between training and deployment environments. Distributionally robust optimization (DRO) provides a principled approach to off-policy evaluation and learning (OPE/L). However, existi…

Cited by 0SourceScholar
2026

Semi-Parametric Contextual Pricing with General Smoothness

ICLR 2026poster

We study the contextual pricing problem, where in each round a seller observes a context, sets a price, and receives a binary purchase signal. We adopt a semi-parametric model in which the demand follows a linear parametric form composed with an unknown link function from a $\beta$-Hölder class. Pri…

Cited by 0SourceScholar
2026

VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild

AAAI 2026technical

Reconstructing topologically consistent facial geometry is crucial for the digital avatar creation pipelines. Existing methods either require tedious manual efforts, lack generalization to in-the-wild data, or are constrained by the limited expressiveness of 3D Morphable Models. To address these lim

Cited by 0SourcePDFScholar
2026

WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Rendering

CVPR 2026

Existing methods achieve high-quality facial albedo capture under controllable lighting, which increases capture cost and limits usability. We propose WildCap, a novel method for high-quality facial albedo capture from a smartphone video recorded in the wild. To disentangle high-quality albedo from

Cited by 0SourcecodeScholar
2025

Improved Confidence Regions and Optimal Algorithms for Online and Offline Linear MNL Bandits

NeurIPS 2025poster

In this work, we consider the data-driven assortment optimization problem under the linear multinomial logit(MNL) choice model. We first establish a improved confidence region for the maximum likelihood estimator (MLE) of the $d$-dimensional linear MNL likelihood function that removes the explicit d…

Cited by 0SourceScholar
2025

Precise Asymptotics and Refined Regret of Variance-Aware UCB

NeurIPS 2025spotlight

In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence Bound (UCB) algorithm that incorporates variance estimates into its decision-making process. More precisely, we provide…

Cited by 0SourceScholar
2024

DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial Purification

NeurIPS 2024poster

Diffusion-based purification has demonstrated impressive robustness as an adversarial defense. However, concerns exist about whether this robustness arises from insufficient evaluation. Our research shows that EOT-based attacks face gradient dilemmas due to global gradient averaging, resulting in in…

Cited by 1SourcePDFScholar
2024

On the Convergence of Projected Bures-Wasserstein Gradient Descent under Euclidean Strong Convexity

ICML 2024poster

The Bures-Wasserstein (BW) gradient descent method has gained considerable attention in various domains, including Gaussian barycenter, matrix recovery and variational inference problems, due to its alignment with the Wasserstein geometry of normal distributions. Despite its popularity, existing con…

Cited by 0SourcePDFScholar
2024

RL in Markov Games with Independent Function Approximation: Improved Sample Complexity Bound under the Local Access Model

AISTATS 2024poster

Efficiently learning equilibria with large state and action spaces in general-sum Markov games while overcoming the curse of multi-agency is a challenging problem. Recent works have attempted to solve this problem by employing independent linear function classes to approximate the marginal $Q$-value…

Cited by 2SourcePDFScholar
2023

Optimal Contextual Bandits with Knapsacks under Realizability via Regression Oracles

AISTATS 2023poster

We study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource consumption in a vector form. The challenge is to maximize the total reward without violating the budget for each resou…

2022

Language-Driven Robot Manipulation With Perspective Disambiguation and Placement Optimization

RA-L 2022

Perspective ambiguity widely exists in human-robot interaction, especially in the language-driven robot manipulation. It is easy to place the object at a wrong position just because of different perspectives. Thus, placement location determination is required in language-driven robot manipulation. I

Cited by 9SourceScholar
2022

Private Streaming SCO in $\ell_p$ geometry with Applications in High Dimensional Online Decision Making

ICML 2022spotlight

Differentially private (DP) stochastic convex optimization (SCO) is ubiquitous in trustworthy machine learning algorithm design. This paper studies the DP-SCO problem with streaming data sampled from a distribution and arrives sequentially. We also consider the continual release model where paramete…

Cited by 16SourcePDFScholar
2021

Disentangled Face Attribute Editing via Instance-Aware Latent Space Search

IJCAI 2021poster

Recent works have shown that a rich set of semantic directions exist in the latent space of Generative Adversarial Networks (GANs), which enables various facial attribute editing applications. However, existing methods may suffer poor attribute variation disentanglement, leading to unwanted change o…

2021

Generalized Linear Bandits with Local Differential Privacy

NeurIPS 2021poster

Contextual bandit algorithms are useful in personalized online decision-making. However, many applications such as personalized medicine and online advertising require the utilization of individual-specific information for effective learning, while user's data should remain private from the server d…