← Search

Nikhil Naik

12 accepted papers

2024

Diffusion Model Alignment Using Direct Preference Optimization

CVPR 2024poster

Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' preferences. In contrast to LLMs human preference learning has not been widely explored in text-to-image diffusion models; th…

Cited by 205SourcePDFScholar
2023

CLIP-Lite: Information Efficient Visual Representation Learning with Language Supervision

AISTATS 2023poster

We propose CLIP-Lite, an information efficient method for visual representation learning by feature alignment with textual annotations. Compared to the previously proposed CLIP model, CLIP-Lite requires only one negative image-text sample pair for every positive image-text sample during the optimiza…

2023

ConRad: Image Constrained Radiance Fields for 3D Generation from a Single Image

NeurIPS 2023poster

We present a novel method for reconstructing 3D objects from a single RGB image. Our method leverages the latest image generation models to infer the hidden 3D structure while remaining faithful to the input image. While existing methods obtain impressive results in generating 3D models from text pr…

Cited by 6SourcePDFScholar
2023

End-to-End Diffusion Latent Optimization Improves Classifier Guidance

ICCV 2023poster

Classifier guidance---using the gradients of an image classifier to steer the generations of a diffusion model---has the potential to dramatically expand the creative control over image generation and editing. However, currently classifier guidance requires either training new noise-aware models to…

Cited by 71PDFcodeScholar
2021

CASTing Your Model: Learning To Localize Improves Self-Supervised Representations

CVPR 2021poster

Recent advances in self-supervised learning (SSL) have largely closed the gap with supervised ImageNet pretraining. Despite their success these methods have been primarily applied to unlabeled ImageNet images, and show marginal gains when trained on larger sets of uncurated images. We hypothesize th…

Cited by 101PDFcodeScholar
2018

Accelerating Neural Architecture Search using Performance Prediction

ICLR 2018workshop

Methods for neural network hyperparameter optimization and meta-modeling are computationally expensive due to the need to train a large number of model configurations. In this paper, we show that standard frequentist regression models can predict the final performance of partially trained model conf…

Cited by 446SourceScholar
2018

Pairwise Confusion for Fine-Grained Visual Classification

ECCV 2018poster

Fine-Grained Visual Classification (FGVC) datasets contain small sample sizes, along with significant intra-class variation and inter-class similarity. While prior work has addressed intra-class variation using localization and segmentation techniques, inter-class similarity may also affect feature…

2017

Designing Neural Network Architectures using Reinforcement Learning

ICLR 2017poster

At present, designing convolutional neural network (CNN) architectures requires both human expertise and labor. New architectures are handcrafted by careful experimentation or modified from a handful of existing networks. We introduce MetaQNN, a meta-modeling algorithm based on reinforcement learnin…

Cited by 1983SourceScholar
2015

A Light Transport Model for Mitigating Multipath Interference in Time-of-Flight Sensors

CVPR 2015poster

Continuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when…

Cited by 105SourcePDFScholar