← Search

Chinmay Hegde

39 accepted papers

2026

OpenThoughts: Data Recipes for Reasoning Models

ICLR 2026oral

Reasoning models have made rapid progress on many benchmarks involving math, code, and science. Yet, there are still many open questions about the best train- ing recipes for reasoning since state-of-the-art models often rely on proprietary datasets with little to no public information available. To…

Cited by 0SourcecodeScholar
2025

Hidden in the Noise: Two-Stage Robust Watermarking for Images

ICLR 2025poster

As the quality of image generators continues to improve, deepfakes become a topic of considerable societal debate. Image watermarking allows responsible model owners to detect and label their AI-generated content, which can mitigate the harm. Yet, current state-of-the-art methods in image watermarki…

2025

LiveBench: A Challenging, Contamination-Limited LLM Benchmark

ICLR 2025spotlight

Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM jud…

2025

VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification

NeurIPS 2025poster

This paper introduces VeriThoughts, a novel dataset designed for reasoning-based Verilog code generation. We establish a new benchmark framework grounded in formal verification methods to evaluate the quality and correctness of generated hardware descriptions. Additionally, we present a suite of spe…

Cited by 0SourcecodeScholar
2025

When Are Concepts Erased From Diffusion Models?

NeurIPS 2025poster

In concept erasure, a model is modified to selectively prevent it from generating a target concept. Despite the rapid development of new methods, it remains unclear how thoroughly these approaches remove the target concept from the model. We begin by proposing two conceptual models for the erasure m…

Cited by 0SourcecodeScholar
2024

BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity

NeurIPS 2024spotlight

We introduce BioTrove, the largest publicly accessible dataset designed to advance AI applications in biodiversity. Curated from the iNaturalist platform and vetted to include only research-grade data, BioTrove contains 161.9 million images, offering unprecedented scale and diversity from three prim…

Cited by 2SourcePDFScholar
2024

Circumventing Concept Erasure Methods For Text-To-Image Generative Models

ICLR 2024poster

Text-to-image generative models can produce photo-realistic images for an extremely broad range of concepts, and their usage has proliferated widely among the general public. On the flip side, these models have numerous drawbacks, including their potential to generate images featuring sexually expli…

2024

DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models

CVPR 2024poster

Recent advances in decentralized deep learning algorithms have demonstrated cutting-edge performance on various tasks with large pre-trained models. However a pivotal prerequisite for achieving this level of competitiveness is the significant communication and computation overheads when updating the…

2024

SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification

NeurIPS 2024poster

Data curation is the problem of how to collect and organize samples into a dataset that supports efficient learning. Despite the centrality of the task, little work has been devoted towards a large-scale, systematic comparison of various curation methods. In this work, we take steps towards a formal…

2024

Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing

NeurIPS 2024poster

G-code (Geometric code) or RS-274 is the most widely used computer numerical control (CNC) and 3D printing programming language. G-code provides machine instructions for the movement of the 3D printer, especially for the nozzle, stage, and extrusion of material for extrusion-based additive manufactu…

Cited by 0SourcePDFScholar
2024

TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks

NeurIPS 2024poster

While tabular classification has traditionally relied on from-scratch training, a recent breakthrough called prior-data fitted networks (PFNs) challenges this approach. Similar to large language models, PFNs make use of pretraining and in-context learning to achieve strong performance on new tasks i…

Cited by 26SourcePDFScholar
2023

Active Learning for Single Neuron Models with Lipschitz Non-Linearities

AISTATS 2023poster

We consider the problem of active learning for single neuron models, also sometimes called “ridge functions”, in the agnostic setting (under adversarial label noise). Such models have been shown to be broadly effective in modeling physical phenomena, and for constructing surrogate data-driven models…

Cited by 12SourcePDFScholar
2023

Implicit Regularization for Group Sparsity

ICLR 2023poster

We study the implicit regularization of gradient descent towards structured sparsity via a novel neural reparameterization, which we call a diagonally grouped linear neural network. We show the following intriguing property of our reparameterization: gradient descent over the squared regression loss…

2022

MDPGT: Momentum-Based Decentralized Policy Gradient Tracking

AAAI 2022technical

We propose a novel policy gradient method for multi-agent reinforcement learning, which leverages two different variance-reduction techniques and does not require large batches over iterations. Specifically, we propose a momentum-based decentralized policy gradient tracking (MDPGT) where a new momen…

2022

Selective Network Linearization for Efficient Private Inference

ICML 2022spotlight

Private inference (PI) enables inferences directly on cryptographically secure data. While promising to address many privacy issues, it has seen limited use due to extreme runtimes. Unlike plaintext inference, where latency is dominated by FLOPs, in PI non-linear functions (namely ReLU) are the bott…

2021

Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data

ICML 2021spotlight

Decentralized learning enables a group of collaborative agents to learn models using a distributed dataset without the need for a central parameter server. Recently, decentralized learning algorithms have demonstrated state-of-the-art results on benchmark data sets, comparable with centralized algor…

2021

Decentralized Deep Learning Using Momentum-Accelerated Consensus

ICASSP 2021accepted

We consider the problem of decentralized deep learning where multiple agents collaborate to learn from a distributed dataset. While several decentralized deep learning approaches exist, the majority consider a central parameter-server topology for aggregating the model parameters from the agents. Ho…

Cited by 0SourceScholar
2021

Differentiable Spline Approximations

NeurIPS 2021poster

The paradigm of differentiable programming has significantly enhanced the scope of machine learning via the judicious use of gradient-based optimization. However, standard differentiable programming methods (such as autodiff) typically require that the machine learning models be differentiable, limi…

2021

Implicit Sparse Regularization: The Impact of Depth and Early Stopping

NeurIPS 2021poster

In this paper, we study the implicit bias of gradient descent for sparse regression. We extend results on regression with quadratic parametrization, which amounts to depth-2 diagonal linear networks, to more general depth-$N$ networks, under more realistic settings of noise and correlated designs. W…

2019

Alternating Phase Projected Gradient Descent with Generative Priors for Solving Compressive Phase Retrieval

ICASSP 2019accepted

The classical problem of phase retrieval arises in various signal acquisition systems. Due to the ill-posed nature of the problem, the solution requires assumptions on the structure of the signal. In the last several years, sparsity and support-based priors have been leveraged successfully to solve…

Cited by 0SourceScholar
2019

Semantic Adversarial Attacks: Parametric Transformations That Fool Deep Classifiers

ICCV 2019poster

Deep neural networks have been shown to exhibit an intriguing vulnerability to adversarial input images corrupted with imperceptible perturbations. However, the majority of adversarial attacks assume global, fine-grained control over the image pixel space. In this paper, we consider a different sett…

Cited by 117PDFcodeScholar
2018

Sub-Diffraction Imaging Using Fourier Ptychography and Structured Sparsity

ICASSP 2018accepted

We consider the problem of super-resolution for sub-diffraction imaging. We adapt conventional Fourier ptychographic approaches, for the case where the images to be acquired have an underlying structured sparsity. We propose some sub-sampling strategies which can be easily adapted to existing ptycho…

Cited by 0SourceScholar
2018

Towards Provable Learning of Polynomial Neural Networks Using Low-Rank Matrix Estimation

AISTATS 2018poster

We study the problem of (provably) learning the weights of a two-layer neural network with quadratic activations. In particular, we focus on the under-parametrized regime where the number of neurons in the hidden layer is (much) smaller than the dimension of the input. Our approach uses a lifting tr…

Cited by 0SourcePDFScholar
2017

Collaborative Deep Learning in Fixed Topology Networks

NeurIPS 2017poster

There is significant recent interest to parallelize deep learning algorithms in order to handle the enormous growth in data and model sizes. While most advances focus on model parallelization and engaging multiple computing agents via using a central parameter server, aspect of data parallelization…

2015

Seismic feature extraction using steiner tree methods

ICASSP 2015accepted

Identifying “interesting” features, such as faults, unconformities, and other events in subsurface images is a challenging task in seismic data processing. Existing state-of-the-art methods usually involve manual intervention in the form of a visual inspection by an expert, but this is time-consumin…

Cited by 0SourceScholar