← Search

Richard Baraniuk

40 accepted papers

2026

Neon: Negative Extrapolation From Self-Training Improves Image Generation

ICLR 2026oral

Scaling generative AI models is bottlenecked by the scarcity of high-quality training data. The ease of synthesizing from a generative model suggests using (unverified) synthetic data to augment a limited corpus of real data for the purpose of fine-tuning in the hope of improving performance. Unfort…

Cited by 0SourcecodeScholar
2025

MITIGATING OVER-EXPLORATION IN LATENT SPACE OPTIMIZATION USING LES

ICML 2025poster

We develop Latent Exploration Score (LES) to mitigate over-exploration in Latent Space Optimization (LSO), a popular method for solving black-box discrete optimization problems. LSO utilizes continuous optimization within the latent space of a Variational Autoencoder (VAE) and is known to be suscept…

Cited by 0SourcePDFScholar
2025

WaLRUS: Wavelets for Long range Representation Using State Space Methods

NeurIPS 2025poster

State-Space Models (SSMs) have proven to be powerful tools for online function approximation and for modeling long-range dependencies in sequential data. While recent methods such as HiPPO have demonstrated strong performance using a few polynomial bases, they remain limited by their reliance on clo…

Cited by 0SourceScholar
2024

Implicit Neural Representations and the Algebra of Complex Wavelets

ICLR 2024poster

Implicit neural representations (INRs) have arisen as useful methods for representing signals on Euclidean domains. By parameterizing an image as a multilayer perceptron (MLP) on Euclidean space, INRs effectively couple spatial and spectral features of the represented signal in a way that is not obv…

Cited by 3SourcePDFScholar
2024

Learning Transferable Features for Implicit Neural Representations

NeurIPS 2024poster

Implicit neural representations (INRs) have demonstrated success in a variety of applications, including inverse problems and neural rendering. An INR is typically trained to capture one signal of interest, resulting in learned neural features that are highly attuned to that signal. Assumed to be le…

Cited by 1SourcePDFScholar
2024

MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education

EMNLP 2024finding

This paper introduces MalAlgoQA, a novel dataset designed to evaluate the counterfactual reasoning capabilities of Large Language Models (LLMs) through a pedagogical approach. The dataset comprises mathematics and reading comprehension questions, each accompanied by four answer choices and their cor…

2024

Pedagogical Alignment of Large Language Models

EMNLP 2024finding

Large Language Models (LLMs), when used in educational settings without pedagogical fine-tuning, often provide immediate answers rather than guiding students through the problem-solving process. This approach falls short of pedagogically best practices and limits their effectiveness as educational t…

2024

Self-Consuming Generative Models Go MAD

ICLR 2024poster

Seismic advances in generative AI algorithms for imagery, text, and other data types have led to the temptation to use AI-synthesized data to train next-generation models. Repeating this process creates an autophagous ("self-consuming") loop whose properties are poorly understood. We conduct a thor…

Cited by 179SourcePDFScholar
2024

Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning

EMNLP 2024finding

The pursuit of personalized education has led to the integration of Large Language Models (LLMs) in developing intelligent tutoring systems. To better understand and adapt to individual student needs, including their misconceptions, LLMs need to be trained on extensive datasets of student-tutor dial…

Cited by 3SourcePDFScholar
2023

A Primal-Dual Framework for Transformers and Neural Networks

ICLR 2023top-25%

Self-attention is key to the remarkable success of transformers in sequence modeling tasks including many applications in natural language processing and computer vision. Like neural network layers, these attention mechanisms are often developed by heuristics and experience. To provide a principled…

Cited by 18SourcePDFScholar
2023

CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science principles

EMNLP 2023long findings

We present a design framework called Conversational Learning with Analytical Step-by-Step Strategies (CLASS) for building advanced Intelligent Tutoring Systems (ITS) powered by high-performance Large Language Models (LLMs). The CLASS framework empowers ITS with two key capabilities. First, through a…

Cited by 0SourcecodeScholar
2023

Mitigating Over-smoothing in Transformers via Regularized Nonlocal Functionals

NeurIPS 2023poster

Transformers have achieved remarkable success in a wide range of natural language processing and computer vision applications. However, the representation capacity of a deep transformer model is degraded due to the over-smoothing issue in which the token representations become identical when the mod…

Cited by 12SourcePDFScholar
2023

Retrieval-based Controllable Molecule Generation

ICLR 2023top-25%

Generating new molecules with specified chemical and biological properties via generative models has emerged as a promising direction for drug discovery. However, existing methods require extensive training/fine-tuning with a large dataset, often unavailable in real-world generation tasks. In this w…

2022

Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent From the Decision Boundary Perspective

CVPR 2022oral

We discuss methods for visualizing neural network decision boundaries and decision regions. We use these visualizations to investigate issues related to reproducibility and generalization in neural network training. We observe that changes in model architecture (and its associate inductive bias) cau…

Cited by 80PDFcodeScholar
2022

Improving Transformers with Probabilistic Attention Keys

ICML 2022spotlight

Multi-head attention is a driving force behind state-of-the-art transformers, which achieve remarkable performance across a variety of natural language processing (NLP) and computer vision tasks. It has been observed that for many applications, those attention heads learn redundant embedding, and mo…

2022

MaGNET: Uniform Sampling from Deep Generative Network Manifolds Without Retraining

ICLR 2022poster

Deep Generative Networks (DGNs) are extensively employed in Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and their variants to approximate the data manifold, and data distribution on that manifold. However, training samples are often obtained based on preferences, costs,…

2022

Open-ended Knowledge Tracing for Computer Science Education

EMNLP 2022main

In educational applications, knowledge tracing refers to the problem of estimating students’ time-varying concept/skill mastery level from their past responses to questions and predicting their future performance.One key limitation of most existing knowledge tracing methods is that they treat studen…

Cited by 50SourcePDFScholar
2022

Parameters or Privacy: A Provable Tradeoff Between Overparameterization and Membership Inference

NeurIPS 2022accept

A surprising phenomenon in modern machine learning is the ability of a highly overparameterized model to generalize well (small error on the test data) even when it is trained to memorize the training data (zero error on the training data). This has led to an arms race towards increasingly overparam…

2022

Polarity Sampling: Quality and Diversity Control of Pre-Trained Generative Networks via Singular Values

CVPR 2022oral

We present Polarity Sampling, a theoretically justified plug-and-play method for controlling the generation quality and diversity of any pre-trained deep generative network (DGN). Leveraging the fact that DGNs are, or can be approximated by, continuous piecewise affine splines, we derive the analyti…

Cited by 38PDFcodeScholar
2021

Math Word Problem Generation with Mathematical Consistency and Problem Context Constraints

EMNLP 2021main

We study the problem of generating arithmetic math word problems (MWPs) given a math equation that specifies the mathematical computation and a context that specifies the problem scenario. Existing approaches are prone to generating MWPs that are either mathematically invalid or have unsatisfactory…

Cited by 52SourcePDFScholar
2021

The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization

NeurIPS 2021poster

Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called “η-t…

2020

Analytical Probability Distributions and Exact Expectation-Maximization for Deep Generative Networks

NeurIPS 2020poster

Deep Generative Networks (DGNs) with probabilistic modeling of their output and latent space are currently trained via Variational Autoencoders (VAEs). In the absence of a known analytical form for the posterior and likelihood expectation, VAEs resort to approximations, including (Amortized) Variati…

Cited by 8SourcePDFScholar
2020

MomentumRNN: Integrating Momentum into Recurrent Neural Networks

NeurIPS 2020poster

Designing deep neural networks is an art that often involves an expensive search over candidate architectures. To overcome this for recurrent neural nets (RNNs), we establish a connection between the hidden state dynamics in an RNN and gradient descent (GD). We then integrate momentum into this fram…

2020

Sub-linear Memory Sketches for Near Neighbor Search on Streaming Data

ICML 2020poster

We present the first sublinear memory sketch that can be queried to find the nearest neighbors in a dataset. Our online sketching algorithm compresses an N element dataset to a sketch of size $O(N^b \log^3 N)$ in $O(N^{(b+1)} \log^3 N)$ time, where $b < 1$. This sketch can correctly report the neare…

Cited by 20SourcePDFScholar
2020

Subspace Fitting Meets Regression: The Effects of Supervision and Orthonormality Constraints on Double Descent of Generalization Errors

ICML 2020poster

We study the linear subspace fitting problem in the overparameterized setting, where the estimated subspace can perfectly interpolate the training examples. Our scope includes the least-squares solutions to subspace fitting tasks with varying levels of supervision in the training data (i.e., the pro…

Cited by 20SourcePDFScholar
2020

The Implicit Regularization of Ordinary Least Squares Ensembles

AISTATS 2020poster

Ensemble methods that average over a collection of independent predictors that are each limited to a subsampling of both the examples and features of the training data command a significant presence in machine learning, such as the ever-popular random forest, yet the nature of the subsampling effect…

2019

Adaptive Estimation for Approximate $k$-Nearest-Neighbor Computations

AISTATS 2019poster

Algorithms often carry out equally many computations for "easy" and "hard" problem instances. In particular, algorithms for finding nearest neighbors typically have the same running time regardless of the particular problem instance. In this paper, we consider the approximate $k$-nearest-neighbor p…

2019

From Hard to Soft: Understanding Deep Network Nonlinearities via Vector Quantization and Statistical Inference

ICLR 2019poster

Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understanding the r\^{o}le played by piecewise affine and convex nonlinearities like the…

Cited by 20SourcePDFScholar
2019

The Geometry of Deep Networks: Power Diagram Subdivision

NeurIPS 2019poster

We study the geometry of deep (neural) networks (DNs) with piecewise affine and convex nonlinearities. The layers of such DNs have been shown to be max-affine spline operators (MASOs) that partition their input space and apply a region-dependent affine mapping to their input to produce their output…

2018

Spline Filters For End-to-End Deep Learning

ICML 2018oral

We propose to tackle the problem of end-to-end learning for raw waveform signals by introducing learnable continuous time-frequency atoms. The derivation of these filters is achieved by defining a functional space with a given smoothness order and boundary conditions. From this space, we derive the…

2018

prDeep: Robust Phase Retrieval with a Flexible Deep Network

ICML 2018oral

Phase retrieval algorithms have become an important component in many modern computational imaging systems. For instance, in the context of ptychography and speckle correlation imaging, they enable imaging past the diffraction limit and through scattering media, respectively. Unfortunately, traditio…

2017

Learned D-AMP: Principled Neural Network based Compressive Image Recovery

NeurIPS 2017poster

Compressive image recovery is a challenging problem that requires fast and accurate algorithms. Recently, neural networks have been applied to this problem with promising results. By exploiting massively parallel GPU processing architectures and oodles of training data, they can run orders of magnit…

2016

Dealbreaker: A Nonlinear Latent Variable Model for Educational Data

ICML 2016poster

Statistical models of student responses on assessment questions, such as those in homeworks and exams, enable educators and computer-based personalized learning systems to gain insights into students’ knowledge using machine learning. Popular student-response models, including the Rasch model and it…

Cited by 11SourcePDFScholar