← Search

Zhaoxian Wu

8 accepted papers

2026

Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory Training

ICML 2026poster

Analog in-memory computing (AIMC) performs computation directly within resistive crossbar arrays, offering an energy-efficient platform to scale large vision and language models. However, non-ideal analog device properties make the training on AIMC devices challenging. In particular, its update asym…

Cited by 0SourceScholar
2025

Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions

NeurIPS 2025oral

As the economic and environmental costs of training and deploying large vision or language models increase dramatically, analog in-memory computing (AIMC) emerges as a promising energy-efficient solution. However, the training perspective, especially its training dynamic, is underexplored. In AIMC h…

Cited by 0SourceScholar
2024

On the Convergence of Single-Timescale Multi-Sequence Stochastic Approximation Without Fixed Point Smoothness

ICASSP 2024accepted

Stochastic approximation (SA) that involves multiple coupled sequences has diverse applications, including but not limited to bilevel optimization, meta learning and reinforcement learning. Unfortunately, the existing multi-timescale analysis of multiple-sequence SA (MSSA) implies a slow convergence…

Cited by 0SourceScholar
2024

Towards Exact Gradient-based Training on Analog In-memory Computing

NeurIPS 2024poster

Given the high economic and environmental costs of using large vision or language models, analog in-memory accelerators present a promising solution for energy-efficient AI. While inference on analog accelerators has been studied recently, the training perspective is underexplored. Recent studies ha…

Cited by 0SourcePDFScholar
2023

Distributed Online Learning With Adversarial Participants In An Adversarial Environment

ICASSP 2023accepted

This paper studies distributed online learning under Byzantine attacks. The performance of an online learning algorithm is characterized by (adversarial) regret, and a sublinear bound is preferred. But we prove that, even with a class of state-of-the-art robust aggregation rules, in an adversarial e…

Cited by 0SourceScholar
2021

Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation

ICASSP 2021accepted

This paper considers the policy evaluation problem in reinforcement learning with agents of a decentralized and directed network. The focus is on decentralized temporal-difference (TD) learning with linear function approximation in the presence of unreliable or even malicious agents, termed as Byzan…

Cited by 0SourceScholar
2020

Resilient to Byzantine Attacks Finite-Sum Optimization Over Networks

ICASSP 2020accepted

This contribution deals with distributed finite-sum optimization for learning over networks in the presence of malicious Byzantine attacks. To cope with such attacks, resilient approaches so far combine stochastic gradient descent (SGD) with different robust aggregation rules. However, the sizeable…

Cited by 0SourceScholar