← Search

Amlan Kar

18 accepted papers

2025

Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions

EMNLP 2025

Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chain-of-thought reasoning—akin to the success observed in language models—via distillation and reinforcement learning. But what about the non-reasoning models already train

Cited by 0SourcePDFScholar
2024

Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata

CVPR 2024highlight

We aim to generate fine-grained 3D geometry from large-scale sparse LiDAR scans abundantly captured by autonomous vehicles (AV). Contrary to prior work on AV scene completion we aim to extrapolate fine geometry from unlabeled and beyond spatial limits of LiDAR scans taking a step towards generating…

Cited by 0SourcePDFScholar
2023

DreamTeacher: Pretraining Image Backbones with Deep Generative Models

ICCV 2023poster

In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well enginee…

Cited by 23PDFScholar
2022

EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

NeurIPS 2022accept

We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need…

2021

ATISS: Autoregressive Transformers for Indoor Scene Synthesis

NeurIPS 2021poster

The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to data synthesis for training and simulation. In this paper, we present ATISS, a novel autoregressive transformer architectur…

2021

Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets

CVPR 2021poster

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a l…

Cited by 37PDFScholar
2020

Interactive Annotation of 3D Object Geometry using 2D Scribbles

ECCV 2020poster

Inferring detailed 3D geometry of the scene is crucial for robotics applications, simulation, and 3D content creation. However, such information is hard to obtain, and thus very few datasets support it. In this paper, we propose an interactive framework for annotating 3D object geometry from both po…

Cited by 17SourcePDFScholar
2020

Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation

ECCV 2020poster

Generation of synthetic data has allowed Machine Learning practitioners to bypass the need for costly collection and labeling of large datasets. Unfortunately the generation of such data often requires experts to carefully design sampling procedures that guarantee creation of realistic scenes. These…

Cited by 109SourcePDFScholar
2019

Meta-Sim: Learning to Generate Synthetic Datasets

ICCV 2019oral

Training models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose Meta-Sim, which learns a generative model of synthetic scenes,…

Cited by 316PDFScholar
2019

Neural Turtle Graphics for Modeling City Road Layouts

ICCV 2019oral

We propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segm…

Cited by 104PDFScholar
2019

Object Instance Annotation With Deep Extreme Level Set Evolution

CVPR 2019poster

In this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved".…

Cited by 98PDFcodeScholar
2018

Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++

CVPR 2018poster

Manually labeling datasets with object masks is extremely time consuming. In this work, we follow the idea of Polygon-RNN to produce polygonal annotations of objects interactively using humans-in-the-loop. We introduce several important improvements to the model: 1) we design a new CNN encoder archi…

Cited by 537SourcePDFScholar
2017

AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos

CVPR 2017poster

We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, fro…

Cited by 201PDFcodeScholar