← Search

Haibo Yang

19 accepted papers

2026

Converge Faster, Talk Less: Hessian-Informed Federated Zeroth-Order Optimization

ICLR 2026poster

Zeroth-order (ZO) optimization enables dimension-free communication in federated learning (FL), making it attractive for fine-tuning of large language models (LLMs) due to significant communication savings. However, existing ZO-FL methods largely overlook curvature information, despite its well-esta…

Cited by 0SourceScholar
2025

Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization

ICLR 2025poster

Federated Learning (FL) offers a promising framework for collaborative and privacy-preserving machine learning across distributed data sources. However, the substantial communication costs associated with FL significantly challenge its efficiency. Specifically, in each communication round, the com…

2025

Exact and Linear Convergence for Federated Learning under Arbitrary Client Participation is Attainable

NeurIPS 2025poster

This work tackles the fundamental challenges in Federated Learning (FL) posed by arbitrary client participation and data heterogeneity, prevalent characteristics in practical FL settings. It is well-established that popular FedAvg-style algorithms struggle with exact convergence and can suffer from…

Cited by 0SourceScholar
2025

FAST: A Lightweight Mechanism Unleashing Arbitrary Client Participation in Federated Learning

IJCAI 2025

Federated Learning (FL) provides a flexible distributed platform where numerous clients with high data and system heterogeneity can collaborate to learn a model. While previous research has shown that FL can handle diverse data, it often completely assumes idealized conditions. In practice, real-wor

Cited by 0SourcePDFScholar
2025

PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective Optimization

AAAI 2025technical

Multi-objective optimization (MOO) lies at the core of many machine learning (ML) applications that involve multiple, potentially conflicting objectives (e.g., multi-task learning, multi-objective reinforcement learning, among many others). Despite the long history of MOO, recent years have witnesse…

2025

STIMULUS: Achieving Fast Convergence and Low Sample Complexity in Stochastic Multi-Objective Learning

UAI 2025

Recently, multi-objective optimization (MOO) has gained attention for its broad applications in ML, operations research, and engineering. However, MOO algorithm design remains in its infancy and many existing MOO methods suffer from unsatisfactory convergence rate and sample complexity performance.

Cited by 0SourcePDFScholar
2025

Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach

NeurIPS 2025poster

Split Federated Learning (SFL) enables scalable training on edge devices by combining the parallelism of Federated Learning (FL) with the computational offloading of Split Learning (SL). Despite its great success, SFL suffers significantly from the well-known straggler issue in distributed learning…

Cited by 0SourceScholar
2024

DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation

ECCV 2024poster

"Learning radiance fields (NeRF) with powerful 2D diffusion models has garnered popularity for text-to-3D generation. Nevertheless, the implicit 3D representations of NeRF lack explicit modeling of meshes and textures over surfaces, and such surface-undefined way may suffer from the issues, e.g., no…

2024

Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning

ICML 2024poster

Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcement learning (MORL) problem and introduces an innovative actor-critic algorithm…

Cited by 4SourcePDFScholar
2024

Understanding Server-Assisted Federated Learning in the Presence of Incomplete Client Participation

ICML 2024poster

Existing works in federated learning (FL) often assume either full client or uniformly distributed client participation. However, in reality, some clients may never participate in FL training (aka incomplete client participation) due to various system heterogeneity factors. A popular solution is the…

Cited by 1SourcePDFScholar
2024

VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation

CVPR 2024poster

Recent innovations on text-to-3D generation have featured Score Distillation Sampling (SDS) which enables the zero-shot learning of implicit 3D models (NeRF) by directly distilling prior knowledge from 2D diffusion models. However current SDS-based models still struggle with intricate text prompts a…

2022

Decentralized Learning for Overparameterized Problems: A Multi-Agent Kernel Approximation Approach

ICLR 2022poster

This work develops a novel framework for communication-efficient distributed learning where the models to be learned are overparameterized. We focus on a class of kernel learning problems (which includes the popular neural tangent kernel (NTK) learning as a special case) and propose a novel {\it mul…

Cited by 0SourcePDFScholar
2022

SAGDA: Achieving $\mathcal{O}(\epsilon^{-2})$ Communication Complexity in Federated Min-Max Learning

NeurIPS 2022accept

Federated min-max learning has received increasing attention in recent years thanks to its wide range of applications in various learning paradigms. Similar to the conventional federated learning for empirical risk minimization problems, communication complexity also emerges as one of the most criti…

Cited by 0SourcePDFScholar
2022

Taming Fat-Tailed (“Heavier-Tailed” with Potentially Infinite Variance) Noise in Federated Learning

NeurIPS 2022accept

In recent years, federated learning (FL) has emerged as an important distributed machine learning paradigm to collaboratively learn a global model with multiple clients, while keeping data local and private. However, a key assumption in most existing works on FL algorithms' convergence analysis is t…

Cited by 13SourcePDFScholar
2021

Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated Learning

ICLR 2021poster

Federated learning (FL) is a distributed machine learning architecture that leverages a large number of workers to jointly learn a model with decentralized data. FL has received increasing attention in recent years thanks to its data privacy protection, communication efficiency and a linear speedup…

Cited by 329SourcePDFScholar
2021

STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning

NeurIPS 2021poster

Federated Learning (FL) refers to the paradigm where multiple worker nodes (WNs) build a joint model by using local data. Despite extensive research, for a generic non-convex FL problem, it is not clear, how to choose the WNs' and the server's update directions, the minibatch sizes, and the local up…

Cited by 73SourcePDFScholar