← Search

Valts Blukis

18 accepted papers

2026

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high-level relationships ("left of," "behind", etc.) but ignore fine-grained spatial unders

Cited by 0SourcecodeScholar
2026

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

CVPR 2026

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can use a wide variety of tools that could augment these capabilities, such as depth e

Cited by 0SourcecodeScholar
2025

3D-MVP: 3D Multiview Pretraining for Manipulation

CVPR 2025poster

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we pro…

Cited by 0SourcePDFScholar
2025

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

CVPR 2025poster

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant chal…

Cited by 9SourcePDFScholar
2024

RVT-2: Learning Precise Manipulation from Few Demonstrations

RSS 2024poster

In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions. To be useful in industrial and household domains, such a system should be capable of learning new tasks with few demonstrations and solving them precisely. Prior works, lik…

2024

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction in Robotics

CoRL 2024poster

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot behavior, VLMs struggle to precisely articulate robot actions usin…

Cited by 53SourcecodeScholar
2023

BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects

CVPR 2023poster

We present a near real-time (10Hz) method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the object. Our method works for arbitrary rigid objects, even when visual texture is largely absent. The object is assu…

2023

CuRobo: Parallelized Collision-Free Robot Motion Generation

ICRA 2023poster

This paper explores the problem of collision-free motion generation for manipulators by formulating it as a global motion optimization problem. We develop a parallel optimization technique to solve this problem and demonstrate its effectiveness on massively parallel GPUs. We show that combining simp…

Cited by 74SourceScholar
2023

ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

ICRA 2023poster

Task planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in nat…

Cited by 893SourcecodeScholar
2023

RVT: Robotic View Transformer for 3D Object Manipulation

CoRL 2023oral

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting scalability. In this work, we propose RVT, a multi-view transform…

Cited by 140SourcecodeScholar
2023

TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation

CVPR 2023poster

Test-time adaptation methods have been gaining attention recently as a practical solution for addressing source-to-target domain gaps by gradually updating the model without requiring labels on the target data. In this paper, we propose a method of test-time adaptation for category-level object pose…

Cited by 39SourcePDFScholar
2022

Correcting Robot Plans with Natural Language Feedback

RSS 2022poster

When humans design cost or goal specifications for robots, they often produce specifications that are ambiguous, under-specified, or beyond planners’ ability to solve. In these cases, corrections provide a valuable tool for human-in-the-loop robot control. Corrections might take the form of new goal…

Cited by 110SourcePDFScholar
2021

A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution

CoRL 2021poster

Natural language provides an accessible and expressive interface to specify long-term tasks for robotic agents. However, non-experts are likely to specify such tasks with high-level instructions, which abstract over specific robot actions through several layers of abstraction. We propose that key to…

Cited by 151SourcecodeScholar
2020

Few-shot Object Grounding and Mapping for Natural Language Robot Instruction Following

CoRL 2020

We study the problem of learning a robot policy to follow natural language instructions that can be easily extended to reason about new objects. We introduce a few-shot language-conditioned object grounding method trained from augmented reality data that uses exemplars to identify objects and align

2019

Learning to Map Natural Language Instructions to Physical Quadcopter Control using Simulated Flight

CoRL 2019

We propose a joint simulation and real-world learning framework for mapping navigation instructions and raw first-person observations to continuous control. Our model estimates the need for environment exploration, predicts the likelihood of visiting environment positions during execution, and contr

2018

Following High-level Navigation Instructions on a Simulated Quadcopter with Imitation Learning

RSS 2018poster

We introduce a method for following high-level navigation instructions by mapping directly from images, instructions and pose estimates to continuous low-level velocity commands for real-time control. The Grounded Semantic Mapping Network (GSMN) is a fully-differentiable neural network architecture…

2018

Mapping Navigation Instructions to Continuous Control Actions with Position-Visitation Prediction

CoRL 2018

We propose an approach for mapping natural language instructions and raw observations to continuous control of a quadcopter drone. Our model predicts interpretable position-visitation distributions indicating where the agent should go during execution and where it should stop, and uses the predicted

2017

Socially competent navigation planning by deep learning of multi-agent path topologies

IROS 2017poster

We present a novel, data-driven framework for planning socially competent robot behaviors in crowded environments. The core of our approach is a topological model of collective navigation behaviors, based on braid groups. This model constitutes the basis for the design of a human-inspired probabilis…

Cited by 56SourceScholar