← Search

Gunnar A Sigurdsson

7 accepted papers

2024

Decision Making for Human-in-the-loop Robotic Agents via Uncertainty-Aware Reinforcement Learning

ICRA 2024poster

In a Human-in-the-Loop paradigm, a robotic agent is able to act mostly autonomously in solving a task, but can request help from an external expert when needed. However, knowing when to request such assistance is critical: too few requests can lead to the robot making mistakes, but too many requests…

Cited by 12SourceScholar
2023

A Simple Approach for Visual Room Rearrangement: 3D Mapping and Semantic Search

ICLR 2023poster

Physically rearranging objects is an important capability for embodied agents. Visual room rearrangement evaluates an agent's ability to rearrange objects in a room to a desired goal based solely on visual input. We propose a simple yet effective method for this problem: (1) search for and map which…

Cited by 4SourcePDFScholar
2023

RREx-BoT: Remote Referring Expressions with a Bag of Tricks

IROS 2023poster

Household robots operate in the same space for years. Such robots incrementally build dynamic maps that can be used for tasks requiring remote object localization. However, benchmarks in robot learning often test generalization through inference on tasks in unobserved environments. In an observed en…

Cited by 9SourceScholar
2020

Visual Grounding in Video for Unsupervised Word Translation

CVPR 2020poster

There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establi…

Cited by 58PDFcodeScholar
2018

Actor and Observer: Joint Modeling of First and Third-Person Videos

CVPR 2018poster

Several theories in cognitive neuroscience suggest that when people interact with the world, or simulate interactions, they do so from a first-person egocentric perspective, and seamlessly transfer knowledge between third-person (observer) and first-person (actor). Despite this, learning such models…

2017

Asynchronous Temporal Fields for Action Recognition

CVPR 2017poster

Actions are more than just movements and trajectories: we cook to eat and we hold a cup to drink from it. A thorough understanding of videos requires going beyond appearance modeling and necessitates reasoning about the sequence of activities, as well as the higher-level constructs such as intention…

Cited by 212PDFcodeScholar
2017

What Actions Are Needed for Understanding Human Actions in Videos?

ICCV 2017poster

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets, evaluation metrics, algorithms, and potential future directions. We…

Cited by 165PDFcodeScholar