← Search

Zhi-Quan Luo

40 accepted papers

2025

Adam-mini: Use Fewer Learning Rates To Gain More

ICLR 2025poster

We propose Adam-mini, an optimizer that achieves on-par or better performance than AdamW with $50$% less memory footprint. Adam-mini reduces memory by cutting down the learning rate resources in Adam (i.e., $1/\sqrt{v}$). By delving into the Hessian structure of neural nets, we find Adam’s $v$ might…

2025

Preserving Diversity in Supervised Fine-Tuning of Large Language Models

ICLR 2025poster

Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However, CE maximizes the likelihood of observed data without accounting for alternative possibilities. As such, CE usually lead…

Cited by 0SourcePDFScholar
2025

ROS: A GNN-based Relax-Optimize-and-Sample Framework for Max-$k$-Cut Problems

ICML 2025poster

The Max-$k$-Cut problem is a fundamental combinatorial optimization challenge that generalizes the classic $\mathcal{NP}$-complete Max-Cut problem. While relaxation techniques are commonly employed to tackle Max-$k$-Cut, they often lack guarantees of equivalence between the solutions of the original…

Cited by 0SourcePDFScholar
2024

Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent

ICML 2024poster

We propose HyperAgent, a reinforcement learning (RL) algorithm based on the hypermodel framework for exploration in RL. HyperAgent allows for the efficient incremental approximation of posteriors associated with an optimal action-value function ($Q^\star$) without the need for conjugacy and follows…

2024

ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) is key to aligning Large Language Models (LLMs), typically paired with the Proximal Policy Optimization (PPO) algorithm. While PPO is a powerful method designed for general reinforcement learning tasks, it is overly sophisticated for LLMs, leading to…

2024

Uniformly Stable Algorithms for Adversarial Training and Beyond

ICML 2024poster

In adversarial machine learning, neural networks suffer from a significant issue known as robust overfitting, where the robust test accuracy decreases over epochs (Rice et al., 2020). Recent research conducted by Xing et al., 2021;Xiao et al., 2022 has focused on studying the uniform stability of ad…

2024

Why Transformers Need Adam: A Hessian Perspective

NeurIPS 2024poster

SGD performs worse than Adam by a significant margin on Transformers, but the reason remains unclear. In this work, we provide an explanation through the lens of Hessian: (i) Transformers are "heterogeneous'': the Hessian spectrum across parameter blocks vary dramatically, a phenomenon we call "bloc…

2023

Imitation Learning from Imperfection: Theoretical Justifications and Algorithms

NeurIPS 2023spotlight

Imitation learning (IL) algorithms excel in acquiring high-quality policies from expert data for sequential decision-making tasks. But, their effectiveness is hampered when faced with limited expert data. To tackle this challenge, a novel framework called (offline) IL with supplementary data has bee…

2023

PAC-Bayesian Spectrally-Normalized Bounds for Adversarially Robust Generalization

NeurIPS 2023poster

Deep neural networks (DNNs) are vulnerable to adversarial attacks. It is found empirically that adversarially robust generalization is crucial in establishing defense algorithms against adversarial attacks. Therefore, it is interesting to study the theoretical guarantee of robust generalization. Thi…

Cited by 11SourcePDFScholar
2023

Provably Efficient Adversarial Imitation Learning with Unknown Transitions

UAI 2023poster

Imitation learning (IL) has proven to be an effective method for learning good policies from expert demonstrations. Adversarial imitation learning (AIL), a subset of IL methods, is particularly promising, but its theoretical foundation in the presence of unknown transitions has yet to be fully devel…

2023

Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks

ICCV 2023poster

Spiking Neural Networks (SNNs) are promising energy-efficient models for neuromorphic computing. For training the non-differentiable SNN models, the backpropagation through time (BPTT) with surrogate gradients (SG) method has achieved high performance. However, this method suffers from considerable…

Cited by 67PDFcodeScholar
2022

Adam Can Converge Without Any Modification On Update Rules

NeurIPS 2022accept

Ever since \citet{reddi2019convergence} pointed out the divergence issue of Adam, many new variants have been designed to obtain convergence. However, vanilla Adam remains exceptionally popular and it works well in practice. Why is there a gap between theory and practice? We point out there is a mis…

Cited by 98SourcePDFScholar
2022

Fast Generic Interaction Detection for Model Interpretability and Compression

ICLR 2022poster

The ability of discovering feature interactions in a black-box model is vital to explainable deep learning. We propose a principled, global interaction detection method by casting our target as a multi-arm bandits problem and solving it swiftly with the UCB algorithm. This adaptive method is free of…

2022

HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning

ICLR 2022poster

Randomized least-square value iteration (RLSVI) is a provably efficient exploration method. However, it is limited to the case where (1) a good feature is known in advance and (2) this feature is fixed during the training. If otherwise, RLSVI suffers an unbearable computational burden to obtain the…

2022

ICASSP-SPGC 2022: Root Cause Analysis for Wireless Network Fault Localization

ICASSP 2022accepted

Localizing the root cause of network faults is crucial to network operation and maintenance (O&M). Significant operational expenses will be saved if the root cause can be identified agilely and accurately. However, this is challenging for human beings due to the complicated wireless environments and…

Cited by 0SourceScholar
2022

Optimal Qos-Aware Network Slicing for Service-Oriented Networks with Flexible Routing

ICASSP 2022accepted

In this paper, we consider the network slicing problem which attempts to map multiple customized virtual network requests (also called services) to a common shared network infrastructure and allocate network resources to meet diverse quality of service (QoS) requirements. We first propose a mixed in…

Cited by 0SourceScholar
2022

Stability Analysis and Generalization Bounds of Adversarial Training

NeurIPS 2022accept

In adversarial machine learning, deep neural networks can fit the adversarial examples on the training dataset but have poor generalization ability on the test set. This phenomenon is called robust overfitting, and it can be observed when adversarially training neural nets on common datasets, includ…

2022

Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation

CVPR 2022poster

Spiking Neural Network (SNN) is a promising energy-efficient AI model when implemented on neuromorphic hardware. However, it is a challenge to efficiently train SNNs due to their non-differentiability. Most existing methods either suffer from high latency (i.e., long simulation time steps), or canno…

Cited by 184PDFcodeScholar
2021

An Efficient Linear Programming Rounding-and-Refinement Algorithm for Large-Scale Network Slicing Problem

ICASSP 2021accepted

In this paper, we consider the network slicing problem which attempts to map multiple customized virtual network requests (also called services) to a common shared network infrastructure and allocate network resources to meet diverse service requirements, and propose an efficient two-stage algorithm…

Cited by 0SourceScholar
2021

Data-Driven Adaptive Network Resource Slicing for Multi-Tenant Networks

ICASSP 2021accepted

Network slicing to support multi-tenancy plays a key role in improving the performance of 5G networks. In this paper, we propose a novel framework for network slicing with the goal of maximizing the expected utilities of tenants in the backhaul and Radio Access Network (RAN), where we reconfigure sl…

Cited by 0SourceScholar
2021

Pushing The Limit of Type I Codebook For Fdd Massive Mimo Beamforming: A Channel Covariance Reconstruction Approach

ICASSP 2021accepted

There is a fundamental trade-off between the channel representation resolution of codebooks and the overheads of feedback communications in the fifth generation new radio (5G NR) frequency division duplex (FDD) massive multiple-input and multiple-output (MIMO) systems. In particular, two types of co…

Cited by 0SourceScholar
2021

When Expressivity Meets Trainability: Fewer than $n$ Neurons Can Work

NeurIPS 2021poster

Modern neural networks are often quite wide, causing large memory and computation costs. It is thus of great interest to train a narrower network. However, training narrow neural nets remains a challenging task. We ask two theoretical questions: Can narrow networks have as strong expressivity as wid…

Cited by 14SourcePDFScholar
2020

A Proximal Dual Consensus Method for Linearly Coupled Multi-Agent Non-Convex Optimization

ICASSP 2020accepted

Motivated by large-scale signal processing and machine learning applications, this paper considers the distributed multi-agent optimization problem for a linearly constrained non-convex problem. Each of the agents owns a local cost function and local variable, but are coupled with each other due to…

Cited by 0SourceScholar
2020

Evaluation of Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Approach for Hearing Devices with Attention Switching

ICASSP 2020accepted

Beamforming is a common technique used to improve speech intelligibility and listening comfort of hearing aids users in a noisy environment. Traditional hearing aids beamforming algorithms require the a priori knowledge of the auditory of the listener, which may not be available in real applications…

Cited by 0SourceScholar
2020

Joint Resource Allocation and Routing for Service Function Chaining with In-Subnetwork Processing

ICASSP 2020accepted

Network Function Virtualization (NFV) is an efficient approach to simplify and accelerate the deployment of diverse network services. A critical challenge lies in mapping Virtual Network Functions (VNFs) to high-volume servers, resource allocation, and traffic routing. In this paper, we study the jo…

Cited by 0SourceScholar
2019

A Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Algorithm for Hearing Devices

ICASSP 2019accepted

Traditional adaptive binaural beamforming algorithms for hearing devices often assume that the target talker is known or can be derived from the listener's look direction. When this assumption is violated, the traditional beamforming algorithms often produce distorted target speech and less than opt…

Cited by 0SourceScholar
2019

A New Quadratic Matrix Inequality Approach to Robust Adaptive Beamforming for General-rank Signal Model

ICASSP 2019accepted

The worst-case robust adaptive beamforming problem for generalrank signal model is considered. This is a nonconvex problem, and an approximate version of it (by introducing a matrix decomposition on the presumed covariance matrix of the desired signal) has been studied in the literature. Herein the…

Cited by 0SourceScholar
2019

Direct Acceleration of SAGA using Sampled Negative Momentum

AISTATS 2019poster

Variance reduction is a simple and effective technique that accelerates convex (or non-convex) stochastic optimization. Among existing variance reduction methods, SVRG and SAGA adopt unbiased gradient estimators and are the most popular variance reduction methods in recent years. Although various ac…

Cited by 62SourcePDFScholar
2019

Scalable Gaussian Process Using Inexact Admm for Big Data

ICASSP 2019accepted

Gaussian process (GP) for machine learning has been well studied over the past two decades and is now widely used in many sectors. However, the design of low-complexity GP models still remains a challenging research problem. In this paper, we propose a novel scalable GP regression model for processi…

Cited by 0SourceScholar
2018

Evaluation of the Penalized Inequality Constrained Minimum Variance Beamformer for Hearing Aids

ICASSP 2018accepted

Beamforming is a common technique used to improve speech intelligibility and listening comfort of hearing aids users in a noisy environment. Traditional beamforming algorithms such as linearly constrained minimum variance (LCMV) beamformer cannot effectively suppress multiple interferences when the…

Cited by 0SourceScholar
2018

Software Defined Resource Allocation for Service-Oriented Networks

ICASSP 2018accepted

To support multiple on-demand services over several fixed communication networks, the network operators must allow flexible customization and fast provision of their network resources. One effective approach is network virtualization, whereby each service is mapped to a virtual subnetwork providing…

Cited by 0SourceScholar
2017

A two-stage optimization approach to the asynchronous multi-sensor registration problem

ICASSP 2017accepted

An important step in multi-sensor data fusion is sensor registration, namely, to estimate sensors' range and azimuth biases from their asynchronous measurements. Assuming the target moves in a straight line with an unknown constant velocity, we propose a two-stage nonlinear least square (LS) approac…

Cited by 0SourceScholar
2017

Comparison of two binaural beamforming approaches for hearing aids

ICASSP 2017accepted

Beamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The bi…

Cited by 0SourceScholar
2017

Traffic engineering for backhaul networks with wireless link scheduling

ICASSP 2017accepted

Traffic engineering (TE) problem is a central component of the next generation cloud-based wireless networks. In this paper, we study a new resource allocation scheme for effective traffic engineering under practical constraints such as the finite buffer size at each node. To reduce the computationa…

Cited by 3SourceScholar
2015

Combining sparse NMF with deep neural network: A new classification-based approach for speech enhancement

ICASSP 2015accepted

In this work, we consider enhancing a target speech from a single-channel noisy observation corrupted by non-stationary noises at low signal-to-noise ratios (SNRs). We take a classification-based approach, where the objective is to estimate an Ideal Binary Mask (IBM) that classifies each time-freque…

Cited by 0SourceScholar
2015

Convergence analysis of alternating direction method of multipliers for a family of nonconvex problems

ICASSP 2015accepted

In this paper, we analyze the behavior of the alternating direction method of multipliers (ADMM), for solving a family of nonconvex problems. Our focus is given to the well-known consensus and sharing problems, both of which have wide applications in signal processing. We show that in the presence o…

Cited by 0SourceScholar
2015

Incorporating spatial information in binaural beamforming for noise suppression in hearing aids

ICASSP 2015accepted

In this paper, we propose a beamforming algorithm for binaural hearing aids with enhanced noise suppression capability. The enhancement is based on incorporating a priori spatial information into the conventional multichannel Wiener filtering (MWF) approach for noise suppression. We develop a low co…

Cited by 11SourceScholar
2015

Semi-asynchronous routing for large scale hierarchical networks

ICASSP 2015accepted

We consider the distributed network routing problem in a large-scale hierarchical network whereby the nodes are partitioned into subnetworks, each managed by a network controller (NC), and there is a central NC to coordinate the operation of the distributed NCs. We propose a semi-asynchronous routin…

Cited by 0SourceScholar