← Search

Robinson Piramuthu

14 accepted papers

2025

Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering

IROS 2025

We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and properties ("What is the color of the car?"), situational queries (such as "Is

Cited by 2SourceScholar
2025

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

EMNLP 2025

Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark g

Cited by 0SourcePDFScholar
2025

T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance Design

ICLR 2025poster

In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a significant advancement by integrating various supervision sig…

Cited by 17SourcePDFScholar
2024

"Don't Forget to Put the Milk Back!" Dataset for Enabling Embodied Agents to Detect Anomalous Situations

RA-L 2024

Home robots intend to make their users lives easier. Our work moves toward more helpful home robots by enabling them to inform their users of dangerous or unsanitary anomalies in the home. Some examples of these anomalies include the user leaving their milk out, forgetting to turn off the stove, or

Cited by 13SourceScholar
2024

Decision Making for Human-in-the-loop Robotic Agents via Uncertainty-Aware Reinforcement Learning

ICRA 2024poster

In a Human-in-the-Loop paradigm, a robotic agent is able to act mostly autonomously in solving a task, but can request help from an external expert when needed. However, knowing when to request such assistance is critical: too few requests can lead to the robot making mistakes, but too many requests…

Cited by 12SourceScholar
2023

A Simple Approach for Visual Room Rearrangement: 3D Mapping and Semantic Search

ICLR 2023poster

Physically rearranging objects is an important capability for embodied agents. Visual room rearrangement evaluates an agent's ability to rearrange objects in a room to a desired goal based solely on visual input. We propose a simple yet effective method for this problem: (1) search for and map which…

Cited by 4SourcePDFScholar
2023

RREx-BoT: Remote Referring Expressions with a Bag of Tricks

IROS 2023poster

Household robots operate in the same space for years. Such robots incrementally build dynamic maps that can be used for tasks requiring remote object localization. However, benchmarks in robot learning often test generalization through inference on tasks in unobserved environments. In an observed en…

Cited by 9SourceScholar
2022

TEACh: Task-Driven Embodied Agents That Chat

AAAI 2022technical

Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to c…

2022

VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

ACL 2022findings

Interactive robots navigating photo-realistic environments need to be trained to effectively leverage and handle the dynamic nature of dialogue in addition to the challenges underlying vision-and-language navigation (VLN). In this paper, we present VISITRON, a multi-modal Transformer-based navigator…

2020

Weakly-Supervised Semantic Segmentation via Sub-Category Exploration

CVPR 2020poster

Existing weakly-supervised semantic segmentation methods using image-level annotations typically rely on initial responses to locate object regions. However, such response maps generated by the classification network usually focus on discriminative object parts, due to the fact that the network does…

Cited by 373PDFcodeScholar
2018

Conditional Image-Text Embedding Networks

ECCV 2018poster

This paper presents an approach for grounding phrases in images which jointly learns multiple text-conditioned embeddings in a single end-to-end model. In order to differentiate text phrases into semantically distinct subspaces, we propose a concept weight branch that automatically assigns phrases t…

2015

ConceptLearner: Discovering Visual Concepts From Weakly Labeled Image Collections

CVPR 2015poster

Discovering visual knowledge from weakly labeled data is crucial to scale up computer vision recognition systems, since it is expensive to obtain fully labeled data for a large number of concept categories. In this paper, we propose ConceptLearner, which is a scalable approach to discover visual con…

Cited by 54SourcePDFScholar
2015

HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition

ICCV 2015poster

In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as…

Cited by 520PDFcodeScholar