← Search

Jonathan Tremblay

30 accepted papers

2026

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high-level relationships ("left of," "behind", etc.) but ignore fine-grained spatial unders

Cited by 0SourcecodeScholar
2026

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

RSS 2026poster

The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often exhibit significant domain overlap between training and ev…

Cited by 0SourceScholar
2026

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

CVPR 2026

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can use a wide variety of tools that could augment these capabilities, such as depth e

Cited by 0SourcecodeScholar
2025

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

CVPR 2025poster

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant chal…

Cited by 9SourcePDFScholar
2024

FactorSim: Generative Simulation via Factorized Representation

NeurIPS 2024poster

Generating simulations to train intelligent agents in game-playing and robotics from natural language input, user input, or task documentation remains an open-ended challenge. Existing approaches focus on parts of this challenge, such as generating reward functions or task hyperparameters. Unlike pr…

Cited by 0SourcePDFScholar
2024

NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows

CVPR 2024poster

We present a method for automatically modifying a NeRF representation based on a single observation of a non-rigid transformed version of the original scene. Our method defines the transformation as a 3D flowspecifically as a weighted linear blending of rigid transformations of 3D anchor points that…

2023

BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects

CVPR 2023poster

We present a near real-time (10Hz) method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the object. Our method works for arbitrary rigid objects, even when visual texture is largely absent. The object is assu…

2023

HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions

IROS 2023poster

We present the HANDAL dataset for category-level object pose estimation and affordance prediction. Unlike previous datasets, ours is focused on robotics-ready manipulable objects that are of the proper size and shape for functional grasping by robot manipulators, such as pliers, utensils, and screwd…

Cited by 31SourcecodeScholar
2023

Parallel Inversion of Neural Radiance Fields for Robust Pose Estimation

ICRA 2023poster

We present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual be…

Cited by 75SourcecodeScholar
2023

ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

ICRA 2023poster

Task planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in nat…

Cited by 893SourcecodeScholar
2023

RGB-Only Reconstruction of Tabletop Scenes for Collision-Free Manipulator Control

ICRA 2023poster

We present a system for collision-free control of a robot manipulator that uses only RGB views of the world. Perceptual input of a tabletop scene is provided by multiple images of an RGB camera (without depth) that is either handheld or mounted on the robot end effector. A NeRF-like process is used…

Cited by 14SourcecodeScholar
2023

TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation

CVPR 2023poster

Test-time adaptation methods have been gaining attention recently as a practical solution for addressing source-to-target domain gaps by gradually updating the model without requiring labels on the target data. In this paper, we propose a method of test-time adaptation for category-level object pose…

Cited by 39SourcePDFScholar
2022

6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark

IROS 2022poster

We present a new dataset for 6-DoF pose estimation of known objects, with a focus on robotic manipulation research. We propose a set of toy grocery objects, whose physical instantiations are readily available for purchase and are appropriately sized for robotic grasping and manipulation. We provide…

Cited by 114SourcecodeScholar
2022

Efficient Geometry-Aware 3D Generative Adversarial Networks

CVPR 2022oral

Unsupervised generation of high-quality multi-view-consistent images and 3D shapes using only collections of single-view 2D photographs has been a long-standing challenge. Existing 3D GANs are either compute-intensive or make approximations that are not 3D-consistent; the former limits quality and r…

Cited by 1564PDFcodeScholar
2022

Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty Estimation

ICRA 2022poster

We propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predi…

Cited by 29SourceScholar
2022

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

CoRL 2022poster

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of thi…

Cited by 157SourcecodeScholar
2022

Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB Image

ICRA 2022poster

Prior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstruct…

Cited by 63SourcecodeScholar
2022

Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated Objects

CVPR 2022poster

Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure, that is, its joints and how they interact with each other. Unf…

Cited by 51PDFcodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Fast Uncertainty Quantification for Deep Object Pose Estimation

ICRA 2021poster

Deep learning-based object pose estimators are often unreliable and overconfident especially when the input image is outside the training domain, for instance, with sim2real transfer. Efficient and robust uncertainty quantification (UQ) in pose estimators is critically needed in many robotic tasks.…

Cited by 35SourceScholar
2021

Hierarchical Planning for Long-Horizon Manipulation with Geometric and Symbolic Scene Graphs

ICRA 2021poster

We present a visually grounded hierarchical planning algorithm for long-horizon manipulation tasks. Our algorithm offers a joint framework of neuro-symbolic task planning and low-level motion generation conditioned on the specified goal. At the core of our approach is a two-level scene graph represe…

Cited by 138SourceScholar
2021

Joint Space Control via Deep Reinforcement Learning

IROS 2021poster

The dominant way to control a robot manipulator uses hand-crafted differential equations leveraging some form of inverse kinematics / dynamics. We propose a simple, versatile joint-level controller that dispenses with differential equations entirely. A deep neural network, trained via model-free rei…

Cited by 25SourceScholar
2021

Multi-view Fusion for Multi-level Robotic Scene Understanding

IROS 2021poster

We present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of…

Cited by 37SourceScholar
2020

Camera-to-Robot Pose Estimation from a Single Image

ICRA 2020poster

We present an approach for estimating the pose of an external camera with respect to a robot using a single RGB image of the robot. The image is processed by a deep neural network to detect 2D projections of keypoints (such as joints) associated with the robot. The network is trained entirely on sim…

Cited by 136SourceScholar
2020

Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning

ICRA 2020poster

Traditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinforcement learning approaches can operate directly from raw sensory inputs with on…

Cited by 75SourceScholar
2020

Indirect Object-to-Robot Pose Estimation from an External Monocular RGB Camera

IROS 2020poster

We present a robotic grasping system that uses a single external monocular RGB camera as input. The object-to-robot pose is computed indirectly by combining the output of two neural networks: one that estimates the object-to-camera pose, and another that estimates the robot-to-camera pose. Both netw…

Cited by 26SourceScholar
2020

Toward Sim-to-Real Directional Semantic Grasping

ICRA 2020poster

We address the problem of directional semantic grasping, that is, grasping a specific object from a specific direction. We approach the problem using deep reinforcement learning via a double deep Q-network (DDQN) that learns to map downsampled RGB input images from a wrist-mounted camera to Q-values…

Cited by 29SourceScholar
2019

PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data

ICCV 2019poster

In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (caused by the dependency of shape and appearance on viewpoint), and 2) small inte…

Cited by 147PDFcodeScholar
2018

Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects

CoRL 2018

Using synthetic data for training deep neural networks for robotic manipulation holds the promise of an almost unlimited amount of pre-labeled training data, generated safely out of harm’s way. One of the key challenges of synthetic data, to date, has been to bridge the so-called reality gap, so tha

2018

Synthetically Trained Neural Networks for Learning Human-Readable Plans from Real-World Demonstrations

ICRA 2018poster

We present a system to infer and execute a human-readable program from a real-world demonstration. The system consists of a series of neural networks to perform perception, program generation, and program execution. Leveraging convolutional pose machines, the perception network reliably detects the…

Cited by 55SourcecodeScholar