← Search

Alexander Toshev

20 accepted papers

2025

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

CVPR 2025poster

We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end,…

Cited by 3SourcePDFScholar
2025

Multimodal Autoregressive Pre-training of Large Vision Encoders

CVPR 2025highlight

We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encode…

2025

UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents

ICCV 2025poster

We build a comprehensive online evaluation benchmark for language-conditioned multi-step task execution on mobile interfaces. Our benchmark strives to evaluate the multi-step planning, reasoning, and visual grounding capabilities of agents, using mobile user interfaces as a concrete testbed. To buil…

Cited by 0SourcePDFScholar
2025

World-consistent Video Diffusion with Explicit 3D Modeling

CVPR 2025highlight

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propo…

Cited by 7SourcePDFScholar
2023

Perceptual Grouping in Contrastive Vision-Language Models

ICCV 2023poster

Recent advances in zero-shot image recognition suggest that vision-language models learn generic visual representations with a high degree of semantic information that may be arbitrarily probed with natural language phrases. Understanding an image, however, is not just about understanding what conte…

Cited by 53PDFScholar
2022

Socially CompliAnt Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation

RA-L 2022

Social navigation is the capability of an autonomous agent, such as a robot, to navigate in a “socially compliant” manner in the presence of other intelligent agents such as humans. With the emergence of autonomously navigating mobile robots in human-populated environments (e.g., domestic service ro

Cited by 195SourceScholar
2021

ReLMoGen: Integrating Motion Generation in Reinforcement Learning for Mobile Manipulation

ICRA 2021poster

Many Reinforcement Learning (RL) approaches use joint control signals (positions, velocities, torques) as action space for continuous control tasks. We propose to lift the action space to a higher level in the form of subgoals for a motion generator (a combination of motion planner and trajectory ex…

Cited by 81SourceScholar
2020

Adversarial Generative Grammars for Human Activity Prediction

ECCV 2020poster

In this paper we propose an adversarial generative grammar model for future prediction. The objective is to learn a model that explicitly captures temporal dependencies, providing a capability to forecast multiple, distinct future activities. Our adversarial grammar is designed so that it can learn…

Cited by 36SourcePDFScholar
2020

Interactive Gibson Benchmark: A Benchmark for Interactive Navigation in Cluttered Environments

RA-L 2020

We present Interactive Gibson Benchmark, the first comprehensive benchmark for training and evaluating Interactive Navigation solutions. Interactive Navigation tasks are robot navigation problems where physical interaction with objects (e.g., pushing) is allowed and even encouraged to reach the goal

Cited by 211SourceScholar
2020

Learning Object-conditioned Exploration using Distributed Soft Actor Critic

CoRL 2020

Object navigation is defined as navigating to an object of a given label in a complex, unexplored environment. In its general form, this problem poses several challenges for Robotics: semantic exploration of unknown environments in search of an object and low-level control. In this work we study obj

Cited by 0SourcePDFScholar
2020

Modeling Long-horizon Tasks as Sequential Interaction Landscapes

CoRL 2020

Task planning over long-time horizons is a challenging and open problem in robotics and its complexity grows exponentially with an increasing number of subtasks. In this paper we present a deep neural network that learns dependencies and transitions across subtasks solely from a set of demonstration

Cited by 0SourcePDFScholar
2019

Long Range Neural Navigation Policies for the Real World

IROS 2019poster

Learned Neural Network based policies have shown promising results for robot navigation. However, most of these approaches fall short of being used on a real robot due to the extensive simulated training they require. These simulations lack the visuals and dynamics of the real world, which makes it…

Cited by 22SourceScholar
2019

Scene Memory Transformer for Embodied Agents in Long-Horizon Tasks

CVPR 2019poster

Many robotic applications require the agent to perform long-horizon tasks in partially observable environments. In such applications, decision making at any step can depend on observations received far in the past. Hence, being able to properly memorize and utilize the long-term history is crucial.…

Cited by 238PDFScholar
2019

Visual Representations for Semantic Target Driven Navigation

ICRA 2019poster

What is a good visual representation for navigation? We study this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a previously unseen environment to a target object, e.g. go to the refrigerator. Instead of acquiring a metric semantic ma…

Cited by 256SourcecodeScholar
2018

Sim2Real Viewpoint Invariant Visual Servoing by Recurrent Control

CVPR 2018poster

Humans are remarkably proficient at controlling their limbs and tools from a wide range of viewpoints. In robotics, this ability is referred to as visual servoing: moving a tool or end-point to a desired location using primarily visual feedback. In this paper, we propose learning viewpoint invariant…

Cited by 134SourcePDFScholar
2017

No Fuss Distance Metric Learning Using Proxies

ICCV 2017poster

We address the problem of distance metric learning (DML), defined as learning a distance consistent with a notion of semantic similarity. Traditionally, for this problem supervision is expressed in the form of sets of points that follow an ordinal relationship -- an anchor point x is similar to a se…

Cited by 827PDFScholar
2017

Towards Accurate Multi-Person Pose Estimation in the Wild

CVPR 2017poster

We propose a method for multi-person detection and 2-D pose estimation that achieves state-of-art results on the challenging COCO keypoints task. It is a simple, yet powerful, top-down approach consisting of two stages. In the first stage, we predict the location and scale of boxes which are likely…

Cited by 1144PDFScholar
2016

Generation and Comprehension of Unambiguous Object Descriptions

CVPR 2016oral

We propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods…

Cited by 1580PDFcodeScholar