← Search

Tianyi Chen

86 accepted papers

2026

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

ICASSP 2026oral

Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is generating pseudo-labels that are both informative and efficient: strong labels, such as those used in HuBERT, improve d…

Cited by 0SourcePDFScholar
2026

CROSS-REPRESENTATION BENCHMARKING IN TIME-SERIES ELECTRONIC HEALTH RECORDS FOR CLINICAL OUTCOME PREDICTION

ICASSP 2026poster

Electronic Health Records (EHRs) enable deep learning for clinical predictions, but the optimal method for representing patient data remains unclear due to inconsistent evaluation practices. We present the first systematic benchmark to compare EHR representation methods, including multivariate time-…

Cited by 0SourcePDFScholar
2026

DRBench: A Realistic Benchmark for Enterprise Deep Research

ICLR 2026poster

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, "What changes should we make to our p…

Cited by 0SourcecodeScholar
2026

Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory Training

ICML 2026poster

Analog in-memory computing (AIMC) performs computation directly within resistive crossbar arrays, offering an energy-efficient platform to scale large vision and language models. However, non-ideal analog device properties make the training on AIMC devices challenging. In particular, its update asym…

Cited by 0SourceScholar
2026

Group-aware Multiscale Ensemble Learning for Test-Time Multimodal Sentiment Analysis

AAAI 2026technical

Multi-modal Sentiment Analysis (MSA) enables machines to perceive human sentiments by integrating multiple modalities such as text, video, and audio. Despite recent progress, most existing methods assume distribution consistency between training and test data—a condition rarely met in real-world sce

Cited by 0SourcePDFScholar
2026

HETEROGENEOUS SELF-SUPERVISED ACOUSTIC PRE-TRAINING WITH LOCAL CONSTRAINTS

ICASSP 2026poster

Self-supervised pre-training using unlabeled data is widely used in automatic speech recognition. In this paper, we propose a new self-supervised pre-training approach to dealing with heterogeneous data. Instead of mixing all the data and minimizing the averaged global loss in the conventional way,…

Cited by 0SourcePDFScholar
2026

Hierarchical Retrieval at Scale: Bridging Transparency and Efficiency

ICML 2026poster

Information retrieval is a core component of many intelligent systems as it enables conditioning of outputs on new and large-scale datasets. While effective, the standard practice of encoding data into high-dimensional representations for similarity search entails large memory and compute footprints…

Cited by 0SourceScholar
2026

Perceptual Flow Network for Visually Grounded Reasoning

ICML 2026poster

Despite the success of LVLMs, general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories, leading to language bias and hallucination. To mitigate this, current methods introduce geometric priors from visual experts as additional supervision. However, we observe that s…

Cited by 0SourceScholar
2026

ProCrop: Learning Aesthetic Image Cropping from Professional Compositions

AAAI 2026technical

Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide c

Cited by 0SourcePDFScholar
2026

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

ICML 2026oral

Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility implies a solution space that strictly supersets the fixed autoregressive trajectory, theoretically unlocking superior re…

Cited by 0SourceScholar
2026

WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference

ICLR 2026poster

The ever-increasing computational demands of large language models (LLMs) make efficient inference a central challenge. While recent advances leverage specialized architectures or selective activation, they typically require (re)training or architectural modifications, limiting their broad applicabi…

Cited by 0SourcecodeScholar
2026

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

ICML 2026poster

Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies a critical reliability gap: state-of-the-art agents face distinct robustness challenges in real-world desktop environmen…

Cited by 0SourceScholar
2025

3DHumanEdit: Multi-modal Body Part-aware Conditioning Information Integration for 3D Human Manipulation

AAAI 2025technical

The rapid advancement of 3D Generative Adversarial Networks (GANs) has significantly enhanced the diversity and quality of generated 3D images. Despite these breakthroughs, the manipulation capabilities of 3D GANs remain unexplored, presenting substantial challenges for practical applications where…

Cited by 0SourcePDFScholar
2025

A First-order Generative Bilevel Optimization Framework for Diffusion Models

ICML 2025poster

Diffusion models, which iteratively denoise data samples to synthesize high-quality outputs, have achieved empirical success across domains. However, optimizing these models for downstream tasks often involves nested bilevel structures, such as tuning hyperparameters for fine-tuning tasks or noise s…

Cited by 0SourcePDFScholar
2025

AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models

CVPR 2025poster

In this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed-form) solutions to the federated learning (FL) with pre-trained models. Our AFL draws inspiration from analytic learning---a gradient-free technique that trains neural networks…

Cited by 2SourcePDFScholar
2025

Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data

NeurIPS 2025spotlight

State Space Models (SSMs) have emerged as promising alternatives to attention mechanisms, with the Mamba architecture demonstrating impressive performance and linear complexity for processing long sequences. However, the fundamental differences between Mamba and Transformer architectures remain inco…

Cited by 0SourceScholar
2025

Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response Functions

NeurIPS 2025oral

As the economic and environmental costs of training and deploying large vision or language models increase dramatically, analog in-memory computing (AIMC) emerges as a promising energy-efficient solution. However, the training perspective, especially its training dynamic, is underexplored. In AIMC h…

Cited by 0SourceScholar
2025

Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression

CVPR 2025poster

Structured pruning and quantization are fundamental techniques used to reduce the size of deep neural networks (DNNs) and typically are applied independently. Applying these techniques jointly via co-optimization has the potential to produce smaller, high-quality models. However, existing joint sche…

2025

Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions

NeurIPS 2025poster

Bilevel optimization, a hierarchical optimization paradigm, has gained significant attention in a wide range of practical applications, notably in the fine-tuning of generative models. However, due to the nested problem structure, most existing algorithms require either the Hessian vector calculatio…

Cited by 0SourceScholar
2025

DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

ICML 2025oral

Despite the success of distillation in large language models (LLMs), most prior work applies identical loss functions to both teacher- and student-generated data. These strategies overlook the synergy between loss formulations and data types, leading to a suboptimal performance boost in student mode…

2025

Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance

ICML 2025spotlight

Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a problem as a semivectorial bilevel optimization problem, whose goal is to optimize a pre-defined preference function, subje…

Cited by 0SourcePDFScholar
2025

Generalizable Humanoid Manipulation with 3D Diffusion Policies

IROS 2025

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensivenes

Cited by 40SourcecodeScholar
2025

Objective Soups: Multilingual Multi-Task Modeling for Speech Processing

NeurIPS 2025poster

The need for training multilingual multi-task speech processing (MSP) models that perform both automatic speech recognition and speech-to-text translation is increasingly evident. However, a significant challenge arises from the conflicts among multiple objectives when using a single model. Multi-ob…

Cited by 0SourceScholar
2025

Optimistic Safety for Online Convex Optimization with Unknown Linear Constraints

AISTATS 2025poster

We study the problem of online convex optimization (OCO) under unknown linear constraints that are either static, or stochastically time-varying. For this problem, we introduce an algorithm that we term Optimistically Safe OCO (OSOCO) and show that it enjoys $\tilde{O}(\sqrt{T})$ regret and no const…

Cited by 0SourceScholar
2025

ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents

NeurIPS 2025poster

Long-horizon tasks requiring multi-step reasoning and dynamic re-planning remain challenging for large language models (LLMs). Sequential prompting methods are prone to context drift, loss of goal information, and recurrent failure cycles, while hierarchical prompting methods often weaken cross-leve…

Cited by 0SourceScholar
2025

SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

ICLR 2025poster

Fine-tuning on task-specific data to boost downstream performance is a crucial step for leveraging Large Language Models (LLMs). However, though fine-tuning enhances the model performance for specialized applications, previous studies have demonstrated that fine-tuning the models on several adversar…

2025

SpotDiff: Spatial Gene Expression Imputation Diffusion with Single-Cell RNA Sequencing Data Integration

AAAI 2025technical

The advent of Spatial Transcriptomics (ST) has revolutionized understanding of tissue architecture by creating high-resolution maps of gene expression patterns. However, the low capture rate of ST leads to significant sparsity. The aim of imputation is to recover biological signals by imputing the d…

Cited by 0SourcePDFScholar
2025

Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering

NeurIPS 2025poster

Existing visual token pruning methods target prompt alignment and visual preservation with static strategies, overlooking the varying relative importance of these objectives across tasks, which leads to inconsistent performance. To address this, we derive the first closed-form error bound for visual…

Cited by 0SourceScholar
2025

WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models

ACL 2025short

Graphical User Interface (GUI) automation relies on accurate GUI grounding. However, obtaining large-scale, high-quality labeled data remains a key challenge, particularly in desktop environments like Windows Operating System (OS). Existing datasets primarily focus on structured web-based elements,…

2024

A Method for Bilevel Optimization with Convex Lower-Level Problem

ICASSP 2024accepted

Gradient-based bilevel optimization methods have been applied to a wide range of applications including hyper-parameter optimization, meta-learning, and model pruning. However, it is known that the bilevel optimization problem is difficult to solve, and the finite-time guarantee has only been establ…

Cited by 0SourceScholar
2024

A Primal-Dual-Assisted Penalty Approach to Bilevel Optimization with Coupled Constraints

NeurIPS 2024poster

Interest in bilevel optimization has grown in recent years, partially due to its relevance for challenging machine-learning problems. Several exciting recent works have been centered around developing efficient gradient-based algorithms that can solve bilevel optimization problems with provable guar…

2024

AttriHuman-3D: Editable 3D Human Avatar Generation with Attribute Decomposition and Indexing

CVPR 2024poster

Editable 3D-aware generation which supports user-interacted editing has witnessed rapid development recently. However existing editable 3D GANs either fail to achieve high-accuracy local editing or suffer from huge computational costs. We propose AttriHuman-3D an editable 3D human generation model w…

Cited by 9SourcePDFScholar
2024

CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering

ECCV 2024poster

"Generalizability and few-shot learning are key challenges in Neural Radiance Fields (NeRF), often due to the lack of a holistic understanding in pixel-level rendering. We introduce CaesarNeRF, an end-to-end approach that leverages scene-level CAlibratEd SemAntic Representation along with pixel-leve…

2024

DREAM: Diffusion Rectification and Estimation-Adaptive Models

CVPR 2024poster

We present DREAM a novel training framework representing Diffusion Rectification and Estimation-Adaptive Models requiring minimal code changes (just three lines) yet significantly enhancing the alignment of training with sampling in diffusion models. DREAM features two components: diffusion rectific…

2024

DistiLLM: Towards Streamlined Distillation for Large Language Models

ICML 2024poster

Knowledge distillation (KD) is widely used for compressing a teacher model to a smaller student model, reducing its inference cost and memory footprint while preserving model capabilities. However, current KD methods for auto-regressive sequence models (e.g., large language models) suffer from missi…

2024

Enhancing In-context Learning via Linear Probe Calibration

AISTATS 2024poster

In-context learning (ICL) is a new paradigm for natural language processing that utilizes Generative Pre-trained Transformer (GPT)-like models. This approach uses prompts that include in-context demonstrations to generate the corresponding output for a new query input. However, applying ICL in real…

2024

FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning

NeurIPS 2024poster

Finding specific preference-guided Pareto solutions that represent different trade-offs among multiple objectives is critical yet challenging in multi-objective problems. Existing methods are restrictive in preference definitions and/or their theoretical guarantees. In this work, we introduce a Fle…

2024

Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization

ICASSP 2024accepted

In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term bi-level joint unsupervised and supervised training (BL-JUST). BL-JUST employs a lower and upper level optimization with an unsupervise…

Cited by 0SourceScholar
2024

ManiFoundation Model for General-Purpose Robotic Manipulation of Contact Synthesis with Arbitrary Objects and Robots

IROS 2024poster

To substantially enhance robot intelligence, there is a pressing need to develop a large model that enables general-purpose robots to proficiently undertake a broad spectrum of manipulation tasks, akin to the versatile task-planning ability exhibited by LLMs. The vast diversity in objects, robots, a…

Cited by 9SourcecodeScholar
2024

Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF

ICML 2024poster

Bilevel optimization has been recently applied to many machine learning tasks. However, their applications have been restricted to the supervised learning setting, where static objective functions with benign structures are considered. But bilevel problems such as incentive design, inverse reinforce…

Cited by 13SourcePDFScholar
2024

RetouchFormer: Semi-supervised High-Quality Face Retouching Transformer with Prior-Based Selective Self-Attention

AAAI 2024technical

Face retouching is to beautify a face image, while preserving the image content as much as possible. It is a promising yet challenging task to remove face imperfections and fill with normal skin. Generic image enhancement methods are hampered by the lack of imperfection localization, which often res…

Cited by 1SourcePDFScholar
2024

SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning

ICML 2024poster

This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward map…

Cited by 2SourcePDFScholar
2024

The Language Model Can Have the Personality: Joint Learning for Personality Enhanced Language Model (Student Abstract)

AAAI 2024technical

With the introduction of large language models, chatbots are becoming more conversational to communicate effectively and capable of handling increasingly complex tasks. To make a chatbot more relatable and engaging, we propose a new language model idea that maps the human-like personality. In this p…

Cited by 1SourcePDFScholar
2024

Towards Exact Gradient-based Training on Analog In-memory Computing

NeurIPS 2024poster

Given the high economic and environmental costs of using large vision or language models, analog in-memory accelerators present a promising solution for energy-efficient AI. While inference on analog accelerators has been studied recently, the training perspective is underexplored. Recent studies ha…

Cited by 0SourcePDFScholar
2024

Variance Reduction Can Improve Trade-Off in Multi-Objective Learning

ICASSP 2024accepted

Many machine learning problems today have multiple objective functions, which are often tackled by the multi-objective learning (MOL) framework. Albeit many encouraging results are obtained by MOL algorithms, a recent theoretical study [1] revealed that these gradient-based MOL methods (e.g., MGDA,…

Cited by 0SourceScholar
2023

Alternating Projected SGD for Equality-constrained Bilevel Optimization

AISTATS 2023poster

Bilevel optimization, which captures the inherent nested structure of machine learning problems, is gaining popularity in many recent applications. Existing works on bilevel optimization mostly consider either the unconstrained problems or the constrained upper-level problems. In this context, this…

2023

An Alternating Optimization Method for Bilevel Problems under the Polyak-Łojasiewicz Condition

NeurIPS 2023poster

Bilevel optimization has recently regained interest owing to its applications in emerging machine learning fields such as hyperparameter optimization, meta-learning, and reinforcement learning. Recent results have shown that simple alternating (implicit) gradient-based algorithms can match the conve…

Cited by 7SourcePDFScholar
2023

Enhancing Robot Program Synthesis Through Environmental Context

NeurIPS 2023poster

Program synthesis aims to automatically generate an executable program that conforms to the given specification. Recent advancements have demonstrated that deep neural methodologies and large-scale pretrained language models are highly proficient in capturing program semantics. For robot programming…

Cited by 3SourcePDFScholar
2023

Exploring Intra-Class Variation Factors With Learnable Cluster Prompts for Semi-Supervised Image Synthesis

CVPR 2023poster

Semi-supervised class-conditional image synthesis is typically performed by inferring and injecting class labels into a conditional Generative Adversarial Network (GAN). The supervision in the form of class identity may be inadequate to model classes with diverse visual appearances. In this paper, w…

Cited by 3SourcePDFScholar
2023

LinSATNet: The Positive Linear Satisfiability Neural Networks

ICML 2023poster

Encoding constraints into neural networks is attractive. This paper studies how to introduce the popular positive linear satisfiability to neural networks. We propose the first differentiable satisfiability layer based on an extension of the classic Sinkhorn algorithm for jointly encoding multiple s…

2023

Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Approach

ICLR 2023top-5%

Many machine learning problems today have multiple objective functions. They appear either in learning with multiple criteria where learning has to make a trade-off between multiple performance metrics such as fairness, safety and accuracy; or, in multi-task learning where multiple tasks are optimiz…

Cited by 55SourcePDFScholar
2023

OTOv2: Automatic, Generic, User-Friendly

ICLR 2023poster

The existing model compression methods via structured pruning typically require complicated multi-stage procedures. Each individual stage necessitates numerous engineering efforts and domain-knowledge from the end-users which prevent their wider applications onto broader scenarios. We propose the se…

2023

Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-Avoidance

NeurIPS 2023poster

Multi-objective learning (MOL) often arises in emerging machine learning problems when multiple learning criteria or tasks need to be addressed. Recent works have developed various _dynamic weighting_ algorithms for MOL, including MGDA and its variants, whose central idea is to find an update direc…

2022

A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled Sequences

NeurIPS 2022accept

Stochastic approximation (SA) with multiple coupled sequences has found broad applications in machine learning such as bilevel learning and reinforcement learning (RL). In this paper, we study the finite-time convergence of nonlinear SA with multiple coupled sequences. Different from existing multi-…

Cited by 19SourcePDFScholar
2022

Is Bayesian Model-Agnostic Meta Learning Better than Model-Agnostic Meta Learning, Provably?

AISTATS 2022poster

Meta learning aims at learning a model that can quickly adapt to unseen tasks. Widely used meta learning methods include model agnostic meta learning (MAML), implicit MAML, Bayesian MAML. Thanks to its ability of modeling uncertainty, Bayesian MAML often has advantageous empirical performance. Howev…

2022

Sharp-MAML: Sharpness-Aware Model-Agnostic Meta Learning

ICML 2022spotlight

Model-agnostic meta learning (MAML) is currently one of the dominating approaches for few-shot meta-learning. Albeit its effectiveness, the optimization of MAML can be challenging due to the innate bilevel problem structure. Specifically, the loss landscape of MAML is much more complex with possibly…

2022

SphericGAN: Semi-Supervised Hyper-Spherical Generative Adversarial Networks for Fine-Grained Image Synthesis

CVPR 2022poster

Generative Adversarial Network (GAN)-based models have greatly facilitated image synthesis. However, the model performance may be degraded when applied to fine-grained data, due to limited training samples and subtle distinction among categories. Different from generic GANs, we address the issue fro…

Cited by 18PDFScholar
2021

An Optimal Stochastic Compositional Optimization Method with Applications to Meta Learning

ICASSP 2021accepted

Stochastic compositional optimization generalizes classic (non-compositional) stochastic optimization to the minimization of com-positions of functions. Each composition may introduce an additional expectation. The series of expectations may be nested. Stochastic compositional optimization is gainin…

Cited by 0SourceScholar
2021

Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation

ICASSP 2021accepted

This paper considers the policy evaluation problem in reinforcement learning with agents of a decentralized and directed network. The focus is on decentralized temporal-difference (TD) learning with linear function approximation in the presence of unreliable or even malicious agents, termed as Byzan…

Cited by 0SourceScholar
2021

CAFE: Catastrophic Data Leakage in Vertical Federated Learning

NeurIPS 2021poster

Recent studies show that private training data can be leaked through the gradients sharing mechanism deployed in distributed machine learning systems, such as federated learning (FL). Increasing batch size to complicate data recovery is often viewed as a promising defense strategy against data leaka…

2021

Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel Problems

NeurIPS 2021spotlight

Stochastic nested optimization, including stochastic compositional, min-max, and bilevel optimization, is gaining popularity in many machine learning applications. While the three problems share a nested structure, existing works often treat them separately, thus developing problem-specific algorit…

Cited by 137SourcePDFScholar
2021

Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement Learning

AAAI 2021technical

This paper deals with distributed reinforcement learning problems with safety constraints. In particular, we consider that a team of agents cooperate in a shared environment, where each agent has its individual reward function and safety constraints that involve all agents' joint actions. As such, t…

Cited by 79SourcePDFScholar
2021

Mask-Embedded Discriminator With Region-Based Semantic Regularization for Semi-Supervised Class-Conditional Image Synthesis

CVPR 2021poster

Semi-supervised generative learning (SSGL) makes use of unlabeled data to achieve a trade-off between the data collection/annotation effort and generation performance, when adequate labeled data are not available. Learning precise class semantics is crucial for class-conditional image synthesis with…

Cited by 6PDFScholar
2021

Only Train Once: A One-Shot Neural Network Training And Pruning Framework

NeurIPS 2021poster

Structured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework t…

2021

Semi-Supervised Single-Stage Controllable GANs for Conditional Fine-Grained Image Generation

ICCV 2021poster

Previous state-of-the-art deep generative models improve fine-grained image generation quality by designing hierarchical model structures and synthesizing images across multiple stages. The learning process is typically performed without any supervision in object categories. To address this issue, w…

Cited by 10PDFScholar
2021

Spatial-Temporal Sequential Hypergraph Network for Crime Prediction with Dynamic Multiplex Relation Learning

IJCAI 2021poster

Crime prediction is crucial for public safety and resource optimization, yet is very challenging due to two aspects: i) the dynamics of criminal patterns across time and space, crime events are distributed unevenly on both spatial and temporal domains; ii) time-evolving dependencies between differen…

2020

Resilient to Byzantine Attacks Finite-Sum Optimization Over Networks

ICASSP 2020accepted

This contribution deals with distributed finite-sum optimization for learning over networks in the presence of malicious Byzantine attacks. To cope with such attacks, resilient approaches so far combine stochastic gradient descent (SGD) with different robust aggregation rules. However, the sizeable…

Cited by 0SourceScholar
2019

Communication-Efficient Distributed Learning via Lazily Aggregated Quantized Gradients

NeurIPS 2019poster

The present paper develops a novel aggregated gradient approach for distributed machine learning that adaptively compresses the gradient communication. The key idea is to first quantize the computed gradients, and then skip less informative quantized gradient communications by reusing outdated gradi…

Cited by 124SourcePDFScholar
2018

LAG: Lazily Aggregated Gradient for Communication-Efficient Distributed Learning

NeurIPS 2018spotlight

This paper presents a new class of gradient methods for distributed machine learning that adaptively skip the gradient calculations to learn with reduced communication and computation. Simple rules are designed to detect slowly-varying gradients and, therefore, trigger the reuse of outdated grad…

Cited by 381SourcePDFScholar
2018

Online Ensemble Multi-kernel Learning Adaptive to Non-stationary and Adversarial Environments

AISTATS 2018poster

Kernel-based methods exhibit well-documented performance in various nonlinear learning tasks. Most of them rely on a preselected kernel, whose prudent choice presumes task-specific prior information. To cope with this limitation, multi-kernel learning has gained popularity thanks to its flexibility…

Cited by 0SourcePDFScholar
2016

Robust geographical load balancing for sustainable data centers

ICASSP 2016accepted

A systematic framework is put forth in this paper to integrate renewable energy sources (RES), distributed storage units, cooling facilities, as well as dynamic pricing into the workload and energy management tasks for a data center network. To cope with RES uncertainty, the resource allocation task…

Cited by 0SourceScholar
2016

Stochastic online control for smart-grid powered MIMO downlink transmissions

ICASSP 2016accepted

An infinite time-horizon resource allocation problem is formulated to maximize the time-averaged multi-input multi-output (MIMO) downlink throughput, subject to a time-averaged energy cost budget. By using the advanced time decoupling technique, a novel stochastic subgradient based online control (S…

Cited by 0SourceScholar