← Search

Navid Azizan

14 accepted papers

2026

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

ICML 2026spotlight

Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge. A fundamental difficulty that distinguishes online RL from standard generative modeling is the lack of direct samples fr…

Cited by 0SourceScholar
2025

Activation-Informed Merging of Large Language Models

NeurIPS 2025poster

Model merging, a method that combines the parameters and embeddings of multiple fine-tuned large language models (LLMs), offers a promising approach to enhance model performance across various tasks while maintaining computational efficiency. This paper introduces Activation-Informed Merging (AIM),…

Cited by 0SourcecodeScholar
2025

Know What You Don't Know: Uncertainty Calibration of Process Reward Models

NeurIPS 2025poster

Process reward models (PRMs) play a central role in guiding inference-time scaling algorithms for large language models (LLMs). However, we observe that even state-of-the-art PRMs can be poorly calibrated. Specifically, they tend to overestimate the success probability that a partial reasoning step…

Cited by 0SourceScholar
2024

Quantifying Representation Reliability in Self-Supervised Learning Models

UAI 2024poster

Self-supervised learning models extract general-purpose representations from data. Quantifying the reliability of these representations is crucial, as many downstream models rely on them as input for their own tasks. To this end, we introduce a formal definition of _representation reliability_: the…

2023

Learning Control-Oriented Dynamical Structure from Data

ICML 2023oral

Even for known nonlinear dynamical systems, feedback controller synthesis is a difficult problem that often requires leveraging the particular structure of the dynamics to induce a stable closed-loop system. For general nonlinear models, including those fit to data, there may not be enough known str…

2023

Online Learning for Traffic Routing under Unknown Preferences

AISTATS 2023poster

In transportation networks, road tolling schemes are a method to cope with the efficiency losses due to selfish user routing, wherein users choose routes to minimize individual travel costs. However, the efficacy of tolling schemes often relies on access to complete information on users’ trip attrib…

2022

A Unified View of SDP-based Neural Network Verification through Completely Positive Programming

AISTATS 2022poster

Verifying that input-output relationships of a neural network conform to prescribed operational specifications is a key enabler towards deploying these networks in safety-critical applications. Semidefinite programming (SDP)-based approaches to Rectified Linear Unit (ReLU) network verification trans…

Cited by 21SourcePDFScholar
2022

Mirror Descent Maximizes Generalized Margin and Can Be Implemented Efficiently

NeurIPS 2022accept

Driven by the empirical success and wide use of deep neural networks, understanding the generalization performance of overparameterized models has become an increasingly popular question. To this end, there has been substantial effort to characterize the implicit bias of the optimization algorithms…

Cited by 27SourcePDFScholar
2021

Adaptive-Control-Oriented Meta-Learning for Nonlinear Systems

RSS 2021poster

Real-time adaptation is imperative to the control of robots operating in complex; dynamic environments. Adaptive control laws can endow even nonlinear systems with good trajectory tracking performance; provided that any uncertain dynamics terms are linearly parameterizable with known nonlinear featu…

2021

Sketching curvature for efficient out-of-distribution detection for deep neural networks

UAI 2021poster

In order to safely deploy Deep Neural Networks (DNNs) within the perception pipelines of real-time decision making systems, there is a need for safeguards that can detect out-of-training-distribution (OoD) inputs both efficiently and accurately. Building on recent work leveraging the local curvature…

Cited by 71SourcePDFScholar
2020

A Study of Generalization of Stochastic Mirror Descent Algorithms on Overparameterized Nonlinear Models

ICASSP 2020accepted

We study the convergence, the implicit regularization and the generalization of stochastic mirror descent (SMD) algorithms in overparameterized nonlinear models, where the number of model parameters exceeds the number of training data points. Due to overpa-rameterization, the training loss has infin…

Cited by 0SourceScholar