← Search

Maneesh Singh

12 accepted papers

2022

AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation

CVPR 2022poster

Powerful priors allow us to perform inference with insufficient information. In this paper, we propose an autoregressive prior for 3D shapes to solve multimodal 3D tasks such as shape completion, reconstruction, and generation. We model the distribution over 3D shapes as a non-sequential autoregress…

Cited by 268PDFcodeScholar
2022

DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection

EMNLP 2022main

We present DocInfer - a novel, end-to-end Document-level Natural Language Inference model that builds a hierarchical document graph enriched through inter-sentence relations (topical, entity-based, concept-based), performs paragraph pruning using the novel SubGraph Pooling layer, followed by optimal…

2022

Hierarchical Semantic Regularization of Latent Spaces in StyleGANs

ECCV 2022poster

"Progress in GANs has enabled the generation of high-resolution photorealistic images of astonishing quality. StyleGANs allow for compelling attribute modification on such images via mathematical operations on the latent style vectors in the W/W+ space that effectively modulate the rich hierarchical…

Cited by 10SourcePDFScholar
2021

Deep Implicit Surface Point Prediction Networks

ICCV 2021poster

Deep neural representations of 3D shapes as implicit functions have been shown to produce high fidelity models surpassing the resolution-memory trade-off faced by the explicit representations using meshes and point clouds. However, most such approaches focus on representing closed shapes. Unsigned d…

Cited by 53PDFScholar
2021

Perturb, Predict & Paraphrase: Semi-Supervised Learning using Noisy Student for Image Captioning

IJCAI 2021poster

Recent semi-supervised learning (SSL) methods are predominantly focused on multi-class classification tasks. Classification tasks allow for easy mixing of class labels during augmentation which does not trivially extend to structured outputs such as word sequences that appear in tasks like image cap…

2020

ProAlignNet: Unsupervised Learning for Progressively Aligning Noisy Contours

CVPR 2020poster

Contour shape alignment is a fundamental but challenging problem in computer vision, especially when the observations are partial, noisy, and largely misaligned. Recent ConvNet-based architectures that were proposed to align image structures tend to fail with contour representation of shapes, mostly…

Cited by 6PDFScholar
2018

Adversarial Learning of Raw Speech Features for Domain Invariant Speech Recognition

ICASSP 2018accepted

Recent advances in neural network based acoustic modelling have shown significant improvements in automatic speech recognition (ASR) performance. In order for acoustic models to be able to handle large acoustic variability, large amounts of labeled data is necessary, which are often expensive to obt…

Cited by 0SourceScholar
2018

Disentangling Factors of Variation with Cycle-Consistent Variational Auto-Encoders

ECCV 2018poster

Generative models that learn disentangled representations for different factors of variation in an image can be very useful for targeted data augmentation. By sampling from the disentangled latent subspace of interest, we can efficiently generate new data necessary for a particular task. Learning di…

Cited by 163SourcePDFScholar
2018

Diverse Image-to-Image Translation via Disentangled Representations

ECCV 2018poster

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for many applications: 1) the lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this work, we present an approach based on disentangled represent…

2017

Unsupervised Representation Learning by Sorting Sequences

ICCV 2017poster

We present an unsupervised representation learning approach using videos without semantic labels. We leverage the temporal coherence as a supervisory signal by formulating representation learning as a sequence sorting task. We take temporally shuffled frames (i.e. in non-chronological order) as inpu…

Cited by 570PDFcodeScholar