← Search

Zhihua Zhang

31 accepted papers

2026

CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal

AAAI 2026technical

Purple flare, a diffuse chromatic aberration artifact commonly found around highlight areas, severely degrades the tone transition and color of the image. Existing traditional methods are based on hand-crafted features, which lack flexibility and rely entirely on fixed priors, while the scarcity of

Cited by 0SourcePDFScholar
2025

A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation

NeurIPS 2025poster

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy $\pi$. Previous…

Cited by 0SourceScholar
2025

Bioinspired Directional Adhesives Enable High Stiffness Layer Jamming in Soft Actuators *

IROS 2025

Soft actuators are inherently flexible and compliant, traits that enhance their adaptability to diverse environments and tasks. However, their low structural stiffness can lead to unpredictable and uncontrollable complex deformations when substantial force is required, thereby compromising their loa

Cited by 0SourceScholar
2025

Follow-the-Perturbed-Leader Nearly Achieves Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

NeurIPS 2025poster

We consider a common case of the combinatorial semi-bandit problem, the $m$-set semi-bandit, where the learner exactly selects $m$ arms from the total $d$ arms. In the adversarial setting, the best regret bound, known to be $\mathcal{O}(\sqrt{nmd})$ for time horizon $n$, is achieved by the well-know…

Cited by 0SourceScholar
2023

A Statistical Analysis of Polyak-Ruppert Averaged Q-Learning

AISTATS 2023poster

We study Q-learning with Polyak-Ruppert averaging (a.k.a., averaged Q-learning) in a discounted markov decision process in synchronous and tabular settings. Under a Lipschitz condition, we establish a functional central limit theorem for the averaged iteration $\bar{\mathbf{Q}}_T$ and show that its…

2023

Diff-Instruct: A Universal Approach for Transferring Knowledge From Pre-trained Diffusion Models

NeurIPS 2023poster

Due to the ease of training, ability to scale, and high sample quality, diffusion models (DMs) have become the preferred option for generative modeling, with numerous pre-trained models available for a wide variety of datasets. Containing intricate information about data distributions, pre-trained D…

2023

Semiparametrically Efficient Off-Policy Evaluation in Linear Markov Decision Processes

ICML 2023poster

We study semiparametrically efficient estimation in off-policy evaluation (OPE) where the underlying Markov decision process (MDP) is linear with a known feature map. We characterize the variance lower bound for regular estimators in the linear MDP setting and propose an efficient estimator whose va…

Cited by 6SourcePDFScholar
2023

Stochastic Distributed Optimization under Average Second-order Similarity: Algorithms and Analysis

NeurIPS 2023poster

We study finite-sum distributed optimization problems involving a master node and $n-1$ local nodes under the popular $\delta$-similarity and $\mu$-strong convexity conditions. We propose two new algorithms, SVRS and AccSVRS, motivated by previous works. The non-accelerated SVRS method combines the…

Cited by 13SourcePDFScholar
2022

Asymptotic Behaviors of Projected Stochastic Approximation: A Jump Diffusion Perspective

NeurIPS 2022accept

In this paper, we consider linearly constrained stochastic approximation problems with federated learning (FL) as a special case. We propose a stochastic approximation algorithm named by LPSA with probabilistic projections to ensure feasibility so that projections are performed with probability $p_n…

Cited by 0SourcePDFScholar
2022

Federated Reinforcement Learning with Environment Heterogeneity

AISTATS 2022poster

We study Federated Reinforcement Learning (FedRL) problem in which $n$ agents collaboratively learn a single policy without sharing the trajectories they collected during agent-environment interaction. In this paper, we stress the constraint of environment heterogeneity, which means $n$ environments…

2022

MR-P: A Parallel Decoding Algorithm for Iterative Refinement Non-Autoregressive Translation

ACL 2022findings

Non-autoregressive translation (NAT) predicts all the target tokens in parallel and significantly speeds up the inference process. The Conditional Masked Language Model (CMLM) is a strong baseline of NAT. It decodes with the Mask-Predict algorithm which iteratively refines the output. Most works abo…

2022

Personalized Federated Learning towards Communication Efficiency, Robustness and Fairness

NeurIPS 2022accept

Personalized Federated Learning faces many challenges such as expensive communication costs, training-time adversarial attacks, and performance unfairness across devices. Recent developments witness a trade-off between a reference model and local models to achieve personalization. We follow the aven…

Cited by 29SourcePDFScholar
2021

Communication-Efficient Distributed SVD via Local Power Iterations

ICML 2021spotlight

We study distributed computing of the truncated singular value decomposition (SVD). We develop an algorithm that we call \texttt{LocalPower} for improving communication efficiency. Specifically, we uniformly partition the dataset among $m$ nodes and alternate between multiple (precisely $p$) local p…

2021

Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data

NeurIPS 2021poster

In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate for (deep) linear networks with spherically symmetric data distribution, which ca…

Cited by 4SourcePDFScholar
2021

Greedy and Random Quasi-Newton Methods with Faster Explicit Superlinear Convergence

NeurIPS 2021poster

In this paper, we follow Rodomanov and Nesterov’s work to study quasi-Newton methods. We focus on the common SR1 and BFGS quasi-Newton methods to establish better explicit (local) superlinear convergence rates. First, based on the greedy quasi-Newton update which greedily selects the direction to ma…

Cited by 19SourcePDFScholar
2020

Lower Complexity Bounds for Finite-Sum Convex-Concave Minimax Optimization Problems

ICML 2020poster

This paper studies the lower bound complexity for minimax optimization problem whose objective function is the average of $n$ individual smooth convex-concave functions. We consider the algorithm which gets access to gradient and proximal oracle for each individual component. For the strongly-convex…

Cited by 24SourcePDFScholar
2019

Lipschitz Generative Adversarial Nets

ICML 2019oral

In this paper we show that generative adversarial networks (GANs) without restriction on the discriminative function space commonly suffer from the problem that the gradient produced by the discriminator is uninformative to guide the generator. By contrast, Wasserstein GAN (WGAN), where the discrimi…

Cited by 107SourcePDFScholar