← Search

Brojeshwar Bhowmick

16 accepted papers

2024

Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments†

ICRA 2024poster

Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-o…

Cited by 10SourceScholar
2024

Task Planning for Object Rearrangement in Multi-Room Environments

AAAI 2024technical

Object rearrangement in a multi-room setup should produce a reasonable plan that reduces the agent's overall travel and the number of steps. Recent state-of-the-art methods fail to produce such plans because they rely on explicit exploration for discovering unseen objects due to partial observabilit…

Cited by 3SourcePDFScholar
2024

Task Planning for Visual Room Rearrangement under Partial Observability

ICLR 2024poster

This paper presents a novel hierarchical task planner under partial observability that empowers an embodied agent to use visual input to efficiently plan a sequence of actions for simultaneous object search and rearrangement in an untidy room, to achieve a desired tidy state. The paper introduces (i…

Cited by 3SourcePDFScholar
2023

Exploring Social Motion Latent Space and Human Awareness for Effective Robot Navigation in Crowded Environments

IROS 2023poster

This work proposes a novel approach to social robot navigation by learning to generate robot controls from a social motion latent space. By leveraging this social motion latent space, the proposed method achieves significant improvements in social navigation metrics such as success rate, navigation…

Cited by 1SourceScholar
2023

Sequence-Agnostic Multi-Object Navigation

ICRA 2023poster

The Multi-Object Navigation (MultiON) task requires a robot to localize an instance (each) of multiple object classes. It is a fundamental task for an assistive robot in a home or a factory. Existing methods for MultiON have viewed this as a direct extension of Object Navigation (ON), the task of lo…

Cited by 10SourceScholar
2022

DoRO: Disambiguation of Referred Object for Embodied Agents

RA-L 2022

Robotic task instructions often involve a referred object that the robot must locate (ground) within the environment. While task intent understanding is an essential part of natural language understanding, less effort is made to resolve ambiguity that may arise while grounding the task. Existing wor

Cited by 19SourceScholar
2022

Emotion-Controllable Generalized Talking Face Generation

IJCAI 2022poster

Despite the significant progress in recent years, very few of the AI-based talking face generation methods attempt to render natural emotions. Moreover, the scope of the methods is majorly limited to the characteristics of the training dataset, hence they fail to generalize to arbitrary unseen faces…

2022

IndoLayout: Leveraging Attention for Extended Indoor Layout Estimation from an RGB Image

IROS 2022poster

In this work, we propose IndoLayout, a novel real-time approach for generating high-quality occupancy maps from an RGB image for indoor scenes. Such occupancy maps are often crucial for path-planning and mapping in indoor environments but are often built using only information contained in the ego v…

Cited by 0SourcecodeScholar
2021

RTVS: A Lightweight Differentiable MPC Framework for Real-Time Visual Servoing

IROS 2021poster

Recent data-driven approaches to visual servoing have shown improved performances over classical methods due to precise feature matching and depth estimation. Some recent servoing approaches use a model predictive control (MPC) framework which generalise well to novel environments and are capable of…

Cited by 6SourceScholar
2020

DeepMPCVS: Deep Model Predictive Control for Visual Servoing

CoRL 2020

The simplicity of the visual servoing approach makes it an attractive option for tasks dealing with vision-based control of robots in many real-world applications. However, attaining precise alignment for unseen environments pose a challenge to existing visual servoing approaches. While classical ap

2020

Simple means Faster: Real-Time Human Motion Forecasting in Monocular First Person Videos on CPU

IROS 2020poster

We present a simple, fast, and light-weight RNN based framework for forecasting future locations of humans in first person monocular videos. The primary motivation for this work was to design a network which could accurately predict future trajectories at a very high rate on a CPU. Typical applicati…

Cited by 6SourceScholar
2020

Speech-driven Facial Animation using Cascaded GANs for Learning of Motion and Texture

ECCV 2020poster

Speech-driven facial animation methods should produce accurate and realistic lip motions with natural expressions and realistic texture portraying target-specific facial characteristics. Moreover, the methods should also be adaptable to any unknown faces and speech quickly during inference. Current…

Cited by 119SourcePDFScholar
2019

Talk to the Vehicle: Language Conditioned Autonomous Navigation of Self Driving Cars

IROS 2019poster

We propose a novel pipeline that blends encodings from natural language and 3D semantic maps obtained from visual imagery to generate local trajectories that are executed by a low-level controller. The pipeline precludes the need for a prior registered map through a local waypoint generator neural n…

Cited by 30SourceScholar
2018

Constructing Category-Specific Models for Monocular Object-SLAM

ICRA 2018poster

We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available. To alleviate the need for huge amounts of labeled data, we…

Cited by 66SourceScholar
2016

Quantification of balance in single limb stance using kinect

ICASSP 2016accepted

This paper presents a novel single limb body balance analysis system which will aid medical practitioners to analyze crucial factor for fall risk minimization, injury prevention, fitness and rehabilitation programs. We use skeleton data obtained from Microsoft Kinect which captures full human body a…

Cited by 0SourceScholar