← Search

Justin Johnson

45 accepted papers

2026

Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation

ICML 2026poster

Latent diffusion models excel at generating high-quality images but lose the benefits of end-to-end modeling. They discard information during image encoding, require a separately trained decoder, and model an auxiliary distribution to the raw data. In this paper, we propose Latent Forcing, a simple …

Cited by 0SourceScholar
2025

Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization

ICCV 2025poster

Since the advent of popular visual generation frameworks like VQGAN and Latent Diffusion Models, state-of-the-art image generation systems have generally been two-stage systems that first tokenize or compress visual data into a lower-dimensional latent space before learning a generative model. Token…

Cited by 0SourcePDFScholar
2025

From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs

ICML 2025poster

3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-order-of-magnitude gap that severely limits performance. We introduce \textbf{\emph{LIFT-GS}}, a practical distillation tech…

2025

Visual Test-time Scaling for GUI Agent Grounding

ICCV 2025poster

We introduce RegionFocus, a visual test-time scaling approach for Vision Language Model Agents. Understanding webpages is challenging due to the visual complexity of GUI images and the large number of interface elements, making accurate action selection difficult. Our approach dynamically zooms in o…

2024

Benchmarking Object Detectors with COCO: A New Path Forward

ECCV 2024poster

"The Common Objects in Context (COCO) dataset has been instrumental in benchmarking object detectors over the past decade. Like every dataset, COCO contains subtle errors and imperfections stemming from its annotation procedure. With the advent of high-performing models, we ask whether these errors…

2024

FAR: Flexible Accurate and Robust 6DoF Relative Camera Pose Estimation

CVPR 2024highlight

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely methods predicting pose directly using neural networks are more robust to limited overl…

Cited by 4SourcePDFScholar
2024

NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

CVPR 2024poster

We address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field…

Cited by 44SourcePDFScholar
2024

PointInfinity: Resolution-Invariant Point Diffusion Models

CVPR 2024poster

We present PointInfinity an efficient family of point cloud diffusion models. Our core idea is to use a transformer-based architecture with a fixed-size resolution-invariant latent representation. This enables efficient training with low-resolution point clouds while allowing high-resolution point c…

Cited by 10SourcePDFScholar
2024

Probing the 3D Awareness of Visual Foundation Models

CVPR 2024poster

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task their intermediate representations are useful for other visual tasks such as detection and segmentation. Given…

2023

Hyperbolic Image-text Representations

ICML 2023poster

Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrasti…

2023

Learning To Predict Scene-Level Implicit 3D From Posed RGBD Data

CVPR 2023poster

We introduce a method that can learn to predict scene-level implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often b…

Cited by 2SourcePDFScholar
2023

Learning Visual Representations via Language-Guided Sampling

CVPR 2023poster

Although an object may appear in numerous contexts, we often describe it in a limited number of ways. Language allows us to abstract away visual variation to represent and communicate concepts. Building on this intuition, we propose an alternative approach to visual representation learning: using la…

2023

Multiview Compressive Coding for 3D Reconstruction

CVPR 2023poster

A central goal of visual recognition is to understand objects and scenes from a single image. 2D recognition has witnessed tremendous progress thanks to large-scale learning and general-purpose representations. But, 3D poses new challenges stemming from occlusions not depicted in the image. Prior wo…

2023

Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

CVPR 2023poster

Recognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D, existing benchmarks are small in size and approaches specia…

2023

Scalable 3D Captioning with Pretrained Models

NeurIPS 2023poster

We introduce Cap3D, an automatic approach for generating descriptive text for 3D objects. This approach utilizes pretrained models from image captioning, image-text alignment, and LLM to consolidate captions from multiple views of a 3D asset, completely side-stepping the time-consuming and costly pr…

2023

Text-To-4D Dynamic Scene Generation

ICML 2023poster

We present MAV3D (Make-A-Video3D), a method for generating three-dimensional dynamic scenes from text descriptions. Our approach uses a 4D dynamic Neural Radiance Field (NeRF), which is optimized for scene appearance, density, and motion consistency by querying a Text-to-Video (T2V) diffusion-based…

2023

Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models

ICCV 2023poster

We present Text2Room, a method for generating room-scale textured 3D meshes from a given text prompt as input. To this end, we leverage pre-trained 2D text-to-image models to synthesize a sequence of images from different poses. In order to lift these outputs into a consistent 3D scene representatio…

Cited by 186PDFcodeScholar
2021

CASTing Your Model: Learning To Localize Improves Self-Supervised Representations

CVPR 2021poster

Recent advances in self-supervised learning (SSL) have largely closed the gap with supervised ImageNet pretraining. Despite their success these methods have been primarily applied to unlabeled ImageNet images, and show marginal gains when trained on larger sets of uncurated images. We hypothesize th…

Cited by 101PDFcodeScholar
2021

PixelSynth: Generating a 3D-Consistent Experience From a Single Image

ICCV 2021poster

Recent advancements in differentiable rendering and 3D reasoning have driven exciting results in novel view synthesis from a single image. Despite realistic results, methods are limited to relatively small view change. In order to synthesize immersive scenes, models must also be able to extrapolate.…

Cited by 88PDFcodeScholar
2021

RedCaps: Web-curated image-text data created by the people, for the people

NeurIPS 2021poster

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text – since web data is noisy, they require complex filtering pipe…

Cited by 183SourcecodeScholar
2021

UnsupervisedR&R: Unsupervised Point Cloud Registration via Differentiable Rendering

CVPR 2021poster

Aligning partial views of a scene into a single whole is essential to understanding one's environment and is a key component of numerous robotics tasks such as SLAM and SfM. Recent approaches have proposed end-to-end systems that can outperform traditional methods by leveraging pose supervision. How…

Cited by 68PDFcodeScholar
2020

SynSin: End-to-End View Synthesis From a Single Image

CVPR 2020oral

View synthesis allows for the generation of new views of a scene given one or more images. This is challenging; it requires comprehensively understanding the 3D scene from images. As a result, current methods typically use multiple images, train on ground-truth depth, or are limited to synthetic dat…

Cited by 510PDFcodeScholar
2019

Mesh R-CNN

ICCV 2019poster

Rapid advances in 2D perception have led to systems that accurately detect objects in real-world images. However, these systems make predictions in 2D, ignoring the 3D structure of the world. Concurrently, advances in 3D shape prediction have mostly focused on synthetic benchmarks and isolated objec…

Cited by 569PDFScholar
2019

On Network Design Spaces for Visual Recognition

ICCV 2019poster

Over the past several years progress in designing better neural network architectures for visual recognition has been substantial. To help sustain this rate of progress, in this work we propose to reexamine the methodology for comparing network architectures. In particular, we introduce a new compar…

Cited by 150PDFcodeScholar
2019

PHYRE: A New Benchmark for Physical Reasoning

NeurIPS 2019poster

Understanding and reasoning about physics is an important ability of intelligent agents. We develop the PHYRE benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment. The benchmark is designed to encourage the development of learning al…

2018

Social GAN: Socially Acceptable Trajectories With Generative Adversarial Networks

CVPR 2018poster

Understanding human motion behavior is critical for autonomous moving platforms (like self-driving cars and social robots) if they are to navigate human-centric environments. This is challenging because human motion is inherently multimodal: given a history of human motion paths, there are many soci…

2017

A Hierarchical Approach for Generating Descriptive Image Paragraphs

CVPR 2017spotlight

Recent progress on image captioning has made it possible to generate novel sentences describing images in natural language, but compressing an image into a single sentence can describe visual content in only coarse detail. While one new captioning approach, dense captioning, can potentially describe…

Cited by 487PDFScholar
2017

CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

CVPR 2017poster

When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover short- comings. Existing benchmarks for visual question answer- ing can help, but have strong biases that models can exploit to correctl…

Cited by 2819PDFScholar
2017

Inferring and Executing Programs for Visual Reasoning

ICCV 2017oral

Existing methods for visual reasoning attempt to directly map inputs to outputs using black-box architectures without explicitly modeling the underlying reasoning processes. As a result, these black-box models often learn to exploit biases in the data rather than learning to perform visual reasoning…

Cited by 677PDFcodeScholar
2016

DenseCap: Fully Convolutional Localization Networks for Dense Captioning

CVPR 2016oral

We introduce the dense captioning task, which requires a computer vision system to both localize and describe salient regions in images in natural language. The dense captioning task generalizes object detection when the descriptions consist of a single word, and Image Captioning when one predicted…

Cited by 1539PDFcodeScholar
2015

Image Retrieval Using Scene Graphs

CVPR 2015poster

This paper develops a novel framework for semantic image retrieval based on the notion of a scene graph. Our scene graphs represent objects ("man", "boat"), attributes of objects ("boat is white") and relationships between objects ("man standing on boat"). We use these scene graphs as queries to ret…

Cited by 1399SourcePDFScholar