← Search

Cong Fang

25 accepted papers

2026

Feasible Fusion: Constrained Joint Estimation under Structural Non-Overlap

ICML 2026poster

Causal inference in modern large-scale systems faces growing challenges, including high-dimensional covariates, multi-valued treatments, massive observational (OBS) data, and limited randomized controlled trial (RCT) samples due to cost constraints. We formalize treatment-induced structural non-over…

Cited by 0SourceScholar
2026

MVR: Multi-view Video Reward Shaping for Reinforcement Learning

ICLR 2026poster

Reward design is of great importance for solving complex tasks with reinforcement learning. Recent studies have explored using image-text similarity produced by vision-language models (VLMs) to augment rewards of a task with visual feedback. A common practice linearly adds VLM scores to task or succ…

Cited by 0SourceScholar
2026

Near-optimal and Efficient First-Order Algorithm for Multi-Task Learning with Shared Linear Representation

ICML 2026poster

Multi-task learning (MTL) has emerged as a pivotal paradigm in machine learning by leveraging shared structures across multiple related tasks. Despite its empirical success, the development of likelihood-based efficiently solvable algorithms—even for shared linear representations—remains largely und…

Cited by 0SourceScholar
2026

PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient Training

ICML 2026poster

Activations have become the primary memory bottleneck in large-batch LLM training. However, existing compression methods fail to exploit structure information of activations, resulting in slow convergence or limited compression. To address this, we bridge the relationship between the algorithm’s fas…

Cited by 0SourceScholar
2026

SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning

ICML 2026poster

In DRL, an agent is trained from a stream of experience. In a continual learning setting, such agents can suffer from \emph{plasticity loss}: their ability to learn new skills from new experiences diminishes over training. Recently, Mixture-of-Experts (MoE) networks have been reported to enable scal…

Cited by 0SourceScholar
2026

What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

ICML 2026poster

Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation (KD) to the emergent phenomenon of Weak-to-Strong (W2S) generalization. While existing studies offer isolated insights, a unified theoretical framewor…

Cited by 0SourceScholar
2025

Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads

NeurIPS 2025poster

Transformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representation, however, often demands substantial memory and compute costs, such as the Key-Value (KV) cache used during auto-regre…

Cited by 0SourcecodeScholar
2025

Learning Curves of Stochastic Gradient Descent in Kernel Regression

ICML 2025poster

This paper considers a canonical problem in kernel regression: how good are the model performances when it is trained by the popular online first-order algorithms, compared to the offline ones, such as ridge and ridgeless regression? In this paper, we analyze the foundational single-pass Stochastic…

Cited by 0SourcePDFScholar
2025

PaZO: Preconditioned Accelerated Zeroth-Order Optimization for Fine-Tuning LLMs

NeurIPS 2025poster

This paper introduces PaZO, a preconditioned accelerated zeroth-order optimization algorithm for fine-tuning large language models (LLMs). First, we theoretically demonstrate the necessity of preconditioning in zeroth-order optimization, proving that zeroth-order stochastic gradient descent (ZO…

Cited by 0SourceScholar
2025

SEPARATE: A Simple Low-rank Projection for Gradient Compression in Modern Large-scale Model Training Process

ICLR 2025poster

Training Large Language Models (LLMs) presents a significant communication bottleneck, predominantly due to the growing scale of the gradient to communicate across multi-device clusters. However, how to mitigate communication overhead in practice remains a formidable challenge due to the weakness of…

Cited by 0SourcePDFScholar
2024

End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations

ICML 2024spotlight

Neuro-symbolic reinforcement learning (NS-RL) has emerged as a promising paradigm for explainable decision-making, characterized by the interpretability of symbolic policies. NS-RL entails structured state representations for tasks with visual observations, but previous methods cannot refine the str…

2024

Optimizing over Multiple Distributions under Generalized Quasar-Convexity Condition

NeurIPS 2024poster

We study a typical optimization model where the optimization variable is composed of multiple probability distributions. Though the model appears frequently in practice, such as for policy problems, it lacks specific analysis in the general setting. For this optimization problem, we propose a new s…

Cited by 0SourcePDFScholar
2024

Quantum Algorithms and Lower Bounds for Finite-Sum Optimization

ICML 2024poster

Finite-sum optimization has wide applications in machine learning, covering important problems such as support vector machines, regression, etc. In this paper, we initiate the study of solving finite-sum optimization problems by quantum computing. Specifically, let $f_1,\ldots,f_n:\mathbb{R}^d\to\ma…

Cited by 4SourcePDFScholar
2024

Relational Learning in Pre-Trained Models: A Theory from Hypergraph Recovery Perspective

ICML 2024poster

Foundation Models (FMs) have demonstrated remarkable insights into the relational dynamics of the world, leading to the crucial question: *how do these models acquire an understanding of world hybrid relations?* Traditional statistical learning, particularly for prediction problems, may overlook the…

Cited by 1SourcePDFScholar
2024

Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics

NeurIPS 2024poster

The deep equilibrium model (DEQ) generalizes the conventional feedforward neural network by fixing the same weights for each layer block and extending the number of layers to infinity. This novel model directly finds the fixed points of such a forward process as features for prediction. Despite emp…

Cited by 0SourcePDFScholar
2024

The Implicit Bias of Heterogeneity towards Invariance: A Study of Multi-Environment Matrix Sensing

NeurIPS 2024poster

Models are expected to engage in invariance learning, which involves distinguishing the core relations that remain consistent across varying environments to ensure the predictions are safe, robust and fair. While existing works consider specific algorithms to realize invariance learning, we show tha…

Cited by 0SourcePDFScholar
2023

Double Randomized Underdamped Langevin with Dimension-Independent Convergence Guarantee

NeurIPS 2023poster

This paper focuses on the high-dimensional sampling of log-concave distributions with composite structures: $p^*(\mathrm{d}x)\propto \exp(-g(x)-f(x))\mathrm{d}x$. We develop a double randomization technique, which leads to a fast underdamped Langevin algorithm with a dimension-independent convergenc…

Cited by 1SourcePDFScholar
2023

Task-Robust Pre-Training for Worst-Case Downstream Adaptation

NeurIPS 2023poster

Pre-training has achieved remarkable success when transferred to downstream tasks. In machine learning, we care about not only the good performance of a model but also its behavior under reasonable shifts of condition. The same philosophy holds when pre-training a foundation model. However, the fou…

Cited by 0SourcePDFScholar
2020

How to Characterize The Landscape of Overparameterized Convolutional Neural Networks

NeurIPS 2020poster

For many initialization schemes, parameters of two randomly initialized deep neural networks (DNNs) can be quite different, but feature distributions of the hidden nodes are similar at each layer. With the help of a new technique called {\it neural network grafting}, we demonstrate that even during…

2020

Improved Analysis of Clipping Algorithms for Non-convex Optimization

NeurIPS 2020poster

Gradient clipping is commonly used in training deep neural networks partly due to its practicability in relieving the exploding gradient problem. Recently, \citet{zhang2019gradient} show that clipped (stochastic) Gradient Descent (GD) converges faster than vanilla GD via introducing a new assumpt…

2019

Complexities in Projection-Free Stochastic Non-convex Minimization

AISTATS 2019poster

For constrained nonconvex minimization problems, we propose a meta stochastic projection-free optimization algorithm, named Normalized Frank Wolfe Updating, that can take any Gradient Estimator (GE) as input. For this algorithm, we prove its convergence rate, regardless of the choice of GE. Using a…

Cited by 35SourcePDFScholar
2019

Learning Compact Partial Differential Equations for Color Images with Efficiency

ICASSP 2019accepted

Learning Partial Differential Equations (LPDEs) from training data for particular tasks has been successfully applied to many image processing problems. In this paper, we aim to learn compact Partial Differential Equations (LCPDEs) for color image tasks by proposing a more effective algorithm. The L…

Cited by 0SourceScholar
2018

SPIDER: Near-Optimal Non-Convex Optimization via Stochastic Path-Integrated Differential Estimator

NeurIPS 2018spotlight

In this paper, we propose a new technique named \textit{Stochastic Path-Integrated Differential EstimatoR} (SPIDER), which can be used to track many deterministic quantities of interests with significantly reduced computational cost. Combining SPIDER with the method of normalized gradient descent,…

Cited by 715SourcePDFScholar
2017

Faster and Non-ergodic O(1/K) Stochastic Alternating Direction Method of Multipliers

NeurIPS 2017poster

We study stochastic convex optimization subjected to linear equality constraints. Traditional Stochastic Alternating Direction Method of Multipliers and its Nesterov's acceleration scheme can only achieve ergodic O(1/\sqrt{K}) convergence rates, where K is the number of iteration. By introducing Var…

Cited by 14SourcePDFScholar