← Search

Enlu Zhou

12 accepted papers

2025

Approximate Bilevel Difference Convex Programming for Bayesian Risk Markov Decision Processes

AAAI 2025technical

We consider infinite-horizon Markov Decision Processes where parameters, such as transition probabilities, are unknown and estimated from data. The popular distributionally robust approach to addressing the parameter uncertainty can sometimes be overly conservative. In this paper, we utilize the rec…

Cited by 0SourcePDFScholar
2023

Cognition Difference-Based Dynamic Trust Network for Distributed Bayesian Data Fusion

IROS 2023poster

Distributed Data Fusion (DDF), as a prevalent technique that empowers scalable, flexible, and robust information fusing, has been employed in various multi-sensor networks operating in uncertain and dynamic environments. This paper proposes a cognition difference-based mechanism to construct a dynam…

Cited by 0SourceScholar
2022

Noise Regularizes Over-parameterized Rank One Matrix Recovery, Provably

AISTATS 2022poster

We investigate the role of noise in optimization algorithms for learning over-parameterized models. Specifically, we consider the recovery of a rank one matrix $Y^*\in R^{d\times d}$ from a noisy observation $Y$ using an over-parameterization model. Specifically, we parameterize the rank one matrix…

Cited by 0SourcePDFScholar
2022

Robust Multi-Objective Bayesian Optimization Under Input Noise

ICML 2022spotlight

Bayesian optimization (BO) is a sample-efficient approach for tuning design parameters to optimize expensive-to-evaluate, black-box performance metrics. In many manufacturing processes, the design parameters are subject to random input noise, resulting in a product that is often less performant than…

2021

Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization

AISTATS 2021poster

Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely unknown. This paper studies this fundamental problem through investigating the nonconvex rectangular matrix factorizatio…

Cited by 14SourcePDFScholar
2019

Toward Understanding the Importance of Noise in Training Neural Networks

ICML 2019oral

Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of deep neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem through training a simple two-layer convolutional neural net…

Cited by 106SourcePDFScholar
2019

Towards Understanding the Importance of Shortcut Connections in Residual Networks

NeurIPS 2019poster

Residual Network (ResNet) is undoubtedly a milestone in deep learning. ResNet is equipped with shortcut connections between layers, and exhibits efficient training using simple first order algorithms. Despite of the great empirical success, the reason behind is far from being well understood. In th…

Cited by 76SourcePDFScholar
2018

Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization

NeurIPS 2018poster

Asynchronous momentum stochastic gradient descent algorithms (Async-MSGD) have been widely used in distributed machine learning, e.g., training large collaborative filtering systems and deep neural networks. Due to current technical limit, however, establishing convergence properties of Async-MSGD f…

Cited by 11SourcePDFScholar