← Search

Tesi Xiao

7 accepted papers

2026

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

ICML 2026poster

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by *projecting future states and selecting goal-directed actions*, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learn…

Cited by 0SourceScholar
2025

COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework

UAI 2025

In LLM alignment and many other ML applications, one often faces the *Multi-Objective Fine-Tuning* (MOFT) problem, *i.e.*, fine-tuning an existing model with datasets labeled w.r.t. different objectives simultaneously. To address the challenge, we propose a *Conditioned One-Shot* fine-tuning framewo

2024

Accelerating Sinkhorn algorithm with sparse Newton iterations

ICLR 2024poster

Computing the optimal transport distance between statistical distributions is a fundamental task in machine learning. One remarkable recent advancement is entropic regularization and the Sinkhorn algorithm, which utilizes only matrix scaling and guarantees an approximated solution with near-linear r…

Cited by 5SourcePDFScholar
2024

Multi-objective Optimization via Wasserstein-Fisher-Rao Gradient Flow

AISTATS 2024poster

Multi-objective optimization (MOO) aims to optimize multiple, possibly conflicting objectives with widespread applications. We introduce a novel interacting particle method for MOO inspired by molecular dynamics simulations. Our approach combines overdamped Langevin and birth-death dynamics, incorpo…

2023

A one-sample decentralized proximal algorithm for non-convex stochastic composite optimization

UAI 2023poster

We focus on decentralized stochastic non-convex optimization, where $n$ agents work together to optimize a composite objective function which is a sum of a smooth term and a non-smooth convex term. To solve this problem, we propose two single-time scale algorithms: \texttt{Prox-DASA} and \texttt{Pro…

2022

A Projection-free Algorithm for Constrained Stochastic Multi-level Composition Optimization

NeurIPS 2022accept

We propose a projection-free conditional gradient-type algorithm for smooth stochastic multi-level composition optimization, where the objective function is a nested composition of $T$ functions and the constraint set is a closed convex set. Our algorithm assumes access to noisy evaluations of the f…

Cited by 8SourcePDFScholar
2020

How Does Noise Help Robustness? Explanation and Exploration under the Neural SDE Framework

CVPR 2020oral

Neural Ordinary Differential Equation (Neural ODE) has been proposed as a continuous approximation to the ResNet architecture. Some commonly used regularization mechanisms in discrete neural networks (e.g., dropout, Gaussian noise) are missing in current Neural ODE networks. In this paper, we propos…

Cited by 70PDFcodeScholar