← Search

Rene Vidal

68 accepted papers

2026

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

ICML 2026poster

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by *projecting future states and selecting goal-directed actions*, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learn…

Cited by 0SourceScholar
2026

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations

ICML 2026poster

Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, motivating the need to find adversarial prompts that realistically elicit such failures. We formulate hallucination elicitation as a constrained optimization problem, where the goal is …

Cited by 0SourceScholar
2026

Shuffling the Data, Extrapolating the Step: Sharper Bias In Constant Step-Size SGD

ICLR 2026poster

From adversarial robustness to multi-agent learning, many machine learning tasks can be cast as finite-sum min–max optimization or, more generally, as variational inequality problems (VIPs). Owing to their simplicity and scalability, stochastic gradient methods with constant step size are widely us…

Cited by 0SourceScholar
2025

A Convex Relaxation Approach to Generalization Analysis for Parallel Positively Homogeneous Networks

AISTATS 2025poster

We propose a general framework for deriving generalization bounds for parallel positively homogeneous neural networks--a class of neural networks whose input-output map decomposes as the sum of positively homogeneous maps. Examples of such networks include matrix factorization and sensing, single…

Cited by 0SourceScholar
2025

Concept Lancet: Image Editing with Compositional Representation Transplant

CVPR 2025poster

Diffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a key challenge: overestimating the edit strength harms visual c…

Cited by 0SourcePDFScholar
2025

Conformal Information Pursuit for Interactively Guiding Large Language Models

NeurIPS 2025poster

A significant use case of instruction-finetuned Large Language Models (LLMs) is to solve question-answering tasks interactively. In this setting, an LLM agent is tasked with making a prediction by sequentially querying relevant information from the user, as opposed to a single-turn conversation. Thi…

Cited by 0SourceScholar
2025

Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares

NeurIPS 2025poster

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the o…

Cited by 0SourceScholar
2025

Disentangling Safe and Unsafe Image Corruptions via Anisotropy and Locality

CVPR 2025poster

State-of-the-art machine learning systems are vulnerable to small perturbations to their input, where _small_ is defined according to a threat model that assigns a positive threat to each perturbation. Most prior works define a task-agnostic, isotropic, and global threat, like the l_p norm, where th…

Cited by 0SourcePDFScholar
2025

Frequency-Guided Posterior Sampling for Diffusion-Based Image Restoration

ICCV 2025poster

Image restoration aims to recover high-quality images from degraded observations. When the degradation process is known, the recovery problem can be formulated as an inverse problem, and in a Bayesian context, the goal is to sample a clean reconstruction given the degraded observation. Recently, mod…

Cited by 0SourcePDFScholar
2025

Guarantees of a Preconditioned Subgradient Algorithm for Overparameterized Asymmetric Low-rank Matrix Recovery

ICML 2025poster

In this paper, we focus on a matrix factorization-based approach for robust recovery of low-rank asymmetric matrices from corrupted measurements. We propose an Overparameterized Preconditioned Subgradient Algorithm (OPSA) and provide, for the first time in the literature, linear convergence rates…

Cited by 2SourcePDFScholar
2025

InCoDe: Interpretable Compressed Descriptions For Image Generation

ICLR 2025poster

Generative models have been successfully applied in diverse domains, from natural language processing to image synthesis. However, despite this success, a key challenge that remains is the ability to control the semantic content of the scene being generated. We argue that adequate control of the gen…

2025

LoRanPAC: Low-rank Random Features and Pre-trained Models for Bridging Theory and Practice in Continual Learning

ICLR 2025poster

The goal of continual learning (CL) is to train a model that can solve multiple tasks presented sequentially. Recent CL approaches have achieved strong performance by leveraging large pre-trained models that generalize well to downstream tasks. However, such methods lack theoretical guarantees, maki…

2025

Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data

NeurIPS 2025poster

Among many mysteries behind the success of deep networks lies the exceptional discriminative power of their learned representations as manifested by the intriguing Neural Collapse (NC) phenomenon, where simple feature structures emerge at the last layer of a trained neural network. Prior works on th…

Cited by 0SourceScholar
2025

SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly deployed in high-risk domains. However, state-of-the-art LLMs often produce hallucinations, raising serious concerns about their reliability. Prior work has explored adversarial attacks for hallucination elicitation in LLMs, but it often produces unreali…

Cited by 0SourcecodeScholar
2025

Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization

AISTATS 2025poster

Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theor…

Cited by 0SourceScholar
2025

Voyaging into Perpetual Dynamic Scenes from a Single View

ICCV 2025poster

The problem of generating a perpetual dynamic scene from a single view is an important problem with widespread applications in augmented and virtual reality, and robotics. However, since dynamic scenes regularly change over time, a key challenge is to ensure that different generated views be consist…

2024

Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification

ICLR 2024poster

Variational Information Pursuit (V-IP) is an interpretable-by-design framework that makes predictions by sequentially selecting a short chain of user-defined, interpretable queries about the data that are most informative for the task. The prediction is based solely on the obtained query answers, wh…

2024

Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization

ICLR 2024poster

This paper studies the problem of training a two-layer ReLU network for binary classification using gradient flow with small initialization. We consider a training dataset with well-separated input vectors: Any pair of input data with the same label are positively correlated, and any pair with diffe…

Cited by 20SourcePDFScholar
2024

Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation

ACL 2024findings

Advancing representation learning in specialized fields like medicine remains challenging due to the scarcity of expert annotations for text and images. To tackle this issue, we present a novel two-stage framework designed to extract high-quality factual statements from free-text radiology reports i…

2024

Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models

ICLR 2024poster

The advent of large pre-trained models has brought about a paradigm shift in both visual representation learning and natural language processing. However, clustering unlabeled images, as a fundamental and classic machine learning problem, still lacks an effective solution, particularly for large-sca…

2024

PaCE: Parsimonious Concept Engineering for Large Language Models

NeurIPS 2024poster

Large Language Models (LLMs) are being used for a wide variety of tasks. While they are capable of generating human-like responses, they can also produce undesirable output including potentially harmful information, racist or sexist language, and hallucinations. Alignment methods are designed to red…

2024

Performance Bounds for Active Binary Testing with Information Maximization

ICML 2024poster

In many applications like experimental design, group testing, and medical diagnosis, the state of a random variable $Y$ is revealed by successively observing the outcomes of binary tests about $Y$. New tests are selected adaptively based on the history of outcomes observed so far. If the number of s…

Cited by 1SourcePDFScholar
2024

Scalable 3D Registration via Truncated Entry-wise Absolute Residuals

CVPR 2024poster

Given an input set of 3D point pairs the goal of outlier-robust 3D registration is to compute some rotation and translation that align as many point pairs as possible. This is an important problem in computer vision for which many highly accurate approaches have been recently proposed. Despite their…

2024

Stochastic Extragradient with Random Reshuffling: Improved Convergence for Variational Inequalities

AISTATS 2024poster

The Stochastic Extragradient (SEG) method is one of the most popular algorithms for solving finite-sum min-max optimization and variational inequality problems (VIPs) appearing in various machine learning tasks. However, existing convergence analyses of SEG focus on its with-replacement variants, wh…

2023

Adversarial Examples Might be Avoidable: The Role of Data Concentration in Adversarial Robustness

NeurIPS 2023poster

The susceptibility of modern machine learning classifiers to adversarial examples has motivated theoretical results suggesting that these might be unavoidable. However, these results can be too general to be applicable to natural data distributions. Indeed, humans are quite robust for tasks involvin…

Cited by 10SourcePDFScholar
2023

Information Maximization Perspective of Orthogonal Matching Pursuit with Applications to Explainable AI

NeurIPS 2023spotlight

Information Pursuit (IP) is a classical active testing algorithm for predicting an output by sequentially and greedily querying the input in order of information gain. However, IP is computationally intensive since it involves estimating mutual information in high-dimensional spaces. This paper expl…

2023

Learning Globally Smooth Functions on Manifolds

ICML 2023poster

Smoothness and low dimensional structures play central roles in improving generalization and stability in learning and statistics. This work combines techniques from semi-infinite constrained learning and manifold regularization to learn representations that are globally smooth on a manifold. To do…

2023

Linear Convergence of Gradient Descent For Finite Width Over-parametrized Linear Networks With General Initialization

AISTATS 2023poster

Recent theoretical analyses of the convergence of gradient descent (GD) to a global minimum for over-parametrized neural networks make strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (spectral, balanced). In this work, we relax these assu…

Cited by 8SourcePDFScholar
2023

Variational Information Pursuit for Interpretable Predictions

ICLR 2023poster

There is a growing interest in the machine learning community in developing predictive algorithms that are interpretable by design. To this end, recent work proposes to sequentially ask interpretable queries about data until a high confidence prediction can be made based on the answers obtained (the…

2022

Global Linear and Local Superlinear Convergence of IRLS for Non-Smooth Robust Regression

NeurIPS 2022accept

We advance both the theory and practice of robust $\ell_p$-quasinorm regression for $p \in (0,1]$ by using novel variants of iteratively reweighted least-squares (IRLS) to solve the underlying non-smooth problem. In the convex case, $p=1$, we prove that this IRLS variant converges globally at a line…

2022

Implicit Bias of Projected Subgradient Method Gives Provable Robust Recovery of Subspaces of Unknown Codimension

ICLR 2022spotlight

Robust subspace recovery (RSR) is the problem of learning a subspace from sample data points corrupted by outliers. Dual Principal Component Pursuit (DPCP) is a robust subspace recovery method that aims to find a basis for the orthogonal complement of the subspace by minimizing the sum of the distan…

Cited by 1SourcePDFScholar
2022

Reverse Engineering $\ell_p$ attacks: A block-sparse optimization approach with recovery guarantees

ICML 2022spotlight

Deep neural network-based classifiers have been shown to be vulnerable to imperceptible perturbations to their input, such as $\ell_p$-bounded norm adversarial attacks. This has motivated the development of many defense methods, which are then broken by new attacks, and so on. This paper focuses on…

Cited by 8SourcePDFScholar
2021

A Nullspace Property for Subspace-Preserving Recovery

ICML 2021spotlight

Much of the theory for classical sparse recovery is based on conditions on the dictionary that are both necessary and sufficient (e.g., nullspace property) or only sufficient (e.g., incoherence and restricted isometry). In contrast, much of the theory for subspace-preserving recovery, the theoretica…

Cited by 4SourcePDFScholar
2021

Dual Principal Component Pursuit for Learning a Union of Hyperplanes: Theory and Algorithms

AISTATS 2021poster

State-of-the-art subspace clustering methods are based on convex formulations whose theoretical guarantees require the subspaces to be low-dimensional. Dual Principal Component Pursuit (DPCP) is a non-convex method that is specifically designed for learning high-dimensional subspaces, such as hyperp…

Cited by 10SourcePDFScholar
2021

Dual Principal Component Pursuit for Robust Subspace Learning: Theory and Algorithms for a Holistic Approach

ICML 2021spotlight

The Dual Principal Component Pursuit (DPCP) method has been proposed to robustly recover a subspace of high-relative dimension from corrupted data. Existing analyses and algorithms of DPCP, however, mainly focus on finding a normal to a single hyperplane that contains the inliers. Although these alg…

Cited by 8SourcePDFScholar
2021

On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear Networks

ICML 2021spotlight

Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon is to study how initialization and overparametrization affe…

Cited by 61SourcePDFScholar
2021

Understanding the Dynamics of Gradient Flow in Overparameterized Linear models

ICML 2021spotlight

We provide a detailed analysis of the dynamics ofthe gradient flow in overparameterized two-layerlinear models. A particularly interesting featureof this model is that its nonlinear dynamics can beexactly solved as a consequence of a large num-ber of conservation laws that constrain the systemto fol…

Cited by 71SourcePDFScholar
2020

A novel variational form of the Schatten-$p$ quasi-norm

NeurIPS 2020poster

The Schatten-$p$ quasi-norm with $p\in(0,1)$ has recently gained considerable attention in various low-rank matrix estimation problems offering significant benefits over relevant convex heuristics such as the nuclear norm. However, due to the nonconvexity of the Schatten-$p$ quasi-norm, minimization…

Cited by 15SourcePDFScholar
2020

Conformal Symplectic and Relativistic Optimization

NeurIPS 2020spotlight

Arguably, the two most popular accelerated or momentum-based optimization methods are Nesterov's accelerated gradient and Polyaks's heavy ball, both corresponding to different discretizations of a particular second order differential equation with a friction term. Such connections with continuous-ti…

2020

Robust Homography Estimation via Dual Principal Component Pursuit

CVPR 2020poster

We revisit robust estimation of homographies over point correspondences between two or three views, a fundamental problem in geometric vision. The analysis serves as a platform to support a rigorous investigation of Dual Principal Component Pursuit (DPCP) as a valid and powerful alternative to RANSA…

Cited by 22PDFScholar
2019

Noisy Dual Principal Component Pursuit

ICML 2019oral

Dual Principal Component Pursuit (DPCP) is a recently proposed non-convex optimization based method for learning subspaces of high relative dimension from noiseless datasets contaminated by as many outliers as the square of the number of inliers. Experimentally, DPCP has proved to be robust to noise…

Cited by 24SourcePDFScholar
2018

Dropout as a Low-Rank Regularizer for Matrix Factorization

AISTATS 2018poster

Regularization for matrix factorization (MF) and approximation problems has been carried out in many different ways. Due to its popularity in deep learning, dropout has been applied also for this class of problems. Despite its solid empirical performance, the theoretical properties of dropout as a r…

Cited by 0SourcePDFScholar
2018

Scalable Exemplar-based Subspace Clustering on Class-Imbalanced Data

ECCV 2018poster

Subspace clustering methods based on expressing each data point as a linear combination of a few other data points (e.g., sparse subspace clustering) have become a popular tool for unsupervised learning due to their empirical success and theoretical guarantees. However, their performance can be affe…

Cited by 99SourcePDFScholar
2017

Provable Self-Representation Based Outlier Detection in a Union of Subspaces

CVPR 2017spotlight

Many computer vision tasks involve processing large amounts of data contaminated by outliers, which need to be detected and rejected. While outlier detection methods based on robust statistics have existed for decades, only recently have methods based on sparse and low-rank representation been devel…

Cited by 141PDFScholar
2017

Temporal Convolutional Networks for Action Segmentation and Detection

CVPR 2017poster

The ability to identify and temporally segment fine-grained human actions throughout a video is crucial for robotics, surveillance, education, and beyond. Typical approaches decouple this problem by first extracting local spatiotemporal features from video frames and then feeding them into a tempora…

Cited by 2196PDFScholar
2016

Oracle Based Active Set Algorithm for Scalable Elastic Net Subspace Clustering

CVPR 2016oral

State-of-the-art subspace clustering methods are based on expressing each data point as a linear combination of other data points while regularizing the matrix of coefficients with l_1, l_2 or nuclear norms. l_1 regularization is guaranteed to give a subspace-preserving affinity (i.e., there are no…

Cited by 318PDFScholar