← Search

Aniket Bera

56 accepted papers

2026

HAVEN: Hierarchical Adversary-Aware Visibility-Enabled Navigation with Cover Utilization Using Deep Transformer Q-Networks

ICRA 2026poster

Autonomous navigation in partially observable environments requires agents to reason beyond immediate sensor input, exploit occlusion, and ensure safety while progressing toward a goal. These challenges arise in many robotics domains, from urban driving and warehouse automation to defense and survei…

2026

Scalable Multi-Robot Informative Path Planning for Target Mapping via Deep Reinforcement Learning

RA-L 2026

Autonomous robots are widely utilized for mapping and exploration tasks due to their cost-effectiveness. Multi-robot systems offer scalability and efficiency, especially in terms of the number of robots deployed in more complex environments. These tasks belong to the set of Multi-Robot Informative P

Cited by 1SourcecodeScholar
2026

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

ICLR 2026poster

Generating interactive 3D scenes from text requires not only synthesizing assets but arranging them with spatial intelligence—support, affordances, and plausibility. However, training data for interactive scenes is dominated by a few indoor datasets, so learning-based methods overfit to in-distribut…

Cited by 0SourceScholar
2026

TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning

ICRA 2026poster

Aerial–ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with aerial semantic context. LiDAR scans are projected into a bird’s-eye-view represe…

2026

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

ICLR 2026poster

Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and game agents require models capable of producing coordinated i…

Cited by 0SourceScholar
2026

Variational Shape Inference for Grasp Diffusion on $\mathrm{SE(3)}$

RA-L 2026

Grasp synthesis is a fundamental task in robotic manipulation which usually has multiple feasible solutions. Multimodal grasp synthesis seeks to generate diverse sets of stable grasps conditioned on object geometry, making the robust learning of geometric features crucial for success. To address thi

Cited by 1SourcecodeScholar
2025

COLLAGE: Collaborative Human-Agent Interaction Generation Using Hierarchical Latent Diffusion and Language Models

ICRA 2025

We propose a novel framework COLLAGE for generating collaborative agent-object-agent interactions by leveraging large language models (LLMs) and hierarchical motion-specific vector-quantized variational autoencoders (VQ-VAEs). Our model addresses the lack of rich datasets in this domain by incorpora

Cited by 3SourcecodeScholar
2025

Dynamic Obstacle Avoidance through Uncertainty-Based Adaptive Planning with Diffusion

IROS 2025

By framing reinforcement learning as a sequence modeling problem, recent work has enabled the use of generative models, such as diffusion models, for planning. While these models are effective in predicting long-horizon state trajectories in deterministic environments, they face challenges in dynami

Cited by 3SourceScholar
2025

EASEIR: Efficient and Adaptive Safe-set Estimation via Implicit Representation for High-dimensional Motion Planning

IROS 2025

Collision-free robotic manipulation is extremely important for all safety-critical applications of robots. Especially for large-scale automation in modern manufacturing facilities where numerous hardware and software systems collaborate in relatively structured environments, accomplishing effectiven

Cited by 0SourceScholar
2025

EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering

IROS 2025

Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches either treat it as static video question answering without active exploration or restrict answers to a closed set of ch

Cited by 11SourcecodeScholar
2025

Go-SLAM: Grounded Object Segmentation and Localization with Gaussian Splatting SLAM

IROS 2025

We introduce Go-Slam, a novel framework that combines 3D Gaussian Splatting SLAM with grounded object segmentation and open-vocabulary querying to enable object-aware 3D scene reconstruction. Go-Slam incrementally builds high-fidelity 3D maps from RGB-D inputs while embedding semantic information by

Cited by 5SourceScholar
2025

Hypergraph-Based Coordinated Task Allocation and Socially-Aware Navigation for Multi-Robot Systems

ICRA 2025

A team of multiple robots seamlessly and safely working in human-filled public environments requires adaptive task allocation and socially-aware navigation that account for dynamic human behavior. Current approaches struggle with highly dynamic pedestrian movement and the need for flexible task allo

Cited by 2SourceScholar
2025

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

ICCV 2025poster

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and de…

Cited by 0SourcePDFScholar
2025

SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models

ICRA 2025

Despite significant advancements in large language models (LLMs) that enhance robot agents' understanding and execution of natural language (NL) commands, ensuring the agents adhere to user-specified constraints remains challenging, particularly for complex commands and long-horizon tasks. To addres

Cited by 20SourcecodeScholar
2025

SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction

CVPR 2025poster

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion triplets, has opened new avenues for training, yielding promising r…

2025

VidSole: A Multimodal Dataset for Joint Kinetics Quantification and Disease Detection with Deep Learning

AAAI 2025technical

Understanding internal joint loading is critical for diagnosing gait-related diseases such as knee osteoarthritis; however, current methods of measuring joint risk factors are time-consuming, expensive, and restricted to lab settings. In this paper, we enable the large-scale, cost-effective biomecha…

Cited by 0SourcePDFScholar
2024

DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

CVPR 2024poster

We have witnessed significant progress in deep learning-based 3D vision ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However existing scene-level datasets for deep learning-based 3D vision limited to either synthetic enviro…

Cited by 85SourcePDFScholar
2024

DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning

AAAI 2024technical

We present DanceAnyWay, a generative learning method to synthesize beat-guided dances of 3D human characters synchronized with music. Our method learns to disentangle the dance movements at the beat frames from the dance movements at all the remaining frames by operating at two hierarchical levels.…

2024

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

CVPR 2024poster

Autonomous driving (AD) has made significant strides in recent years. However existing frameworks struggle to interpret and execute spontaneous user instructions such as "overtake the car ahead." Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to br…

2024

Optimizing Crowd-Aware Multi-Agent Path Finding through Local Communication with Graph Neural Networks

IROS 2024poster

Multi-Agent Path Finding (MAPF) in crowded environments presents a challenging problem in motion planning, aiming to find collision-free paths for all agents in the system. MAPF finds a wide range of applications in various domains, including aerial swarms, autonomous warehouse robotics, and self-dr…

Cited by 1SourceScholar
2024

Quantifying Uncertainty in Motion Prediction with Variational Bayesian Mixture

CVPR 2024poster

Safety and robustness are crucial factors in developing trustworthy autonomous vehicles. One essential aspect of addressing these factors is to equip vehicles with the capability to predict future trajectories for all moving objects in the surroundings and quantify prediction uncertainties. In this…

2024

Scaling Safe Multi-Agent Control for Signal Temporal Logic Specifications

CoRL 2024poster

Existing methods for safe multi-agent control using logic specifications like Signal Temporal Logic (STL) often face scalability issues. This is because they rely either on single-agent perspectives or on Mixed Integer Linear Programming (MILP)-based planners, which are complex to optimize. These me…

Cited by 1SourcecodeScholar
2024

TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing

ACL 2024findings

Given a source and its edited version performed based on human instructions in natural language, how do we extract the underlying edit operations, to automatically replicate similar edits on other images? This is the problem of reverse designing, and we present TAME-RD, a model to solve this problem…

Cited by 0SourcePDFScholar
2024

Trajectory Prediction for Robot Navigation using Flow-Guided Markov Neural Operator

ICRA 2024poster

Predicting pedestrian movements remains a complex and persistent challenge in robot navigation research. We must evaluate several factors to achieve accurate predictions, such as pedestrian interactions, the environment, crowd density, and social and cultural norms. Accurate prediction of pedestrian…

Cited by 5SourceScholar
2024

TrustNavGPT: Modeling Uncertainty to Improve Trustworthiness of Audio-Guided LLM-Based Robot Navigation

IROS 2024poster

Large language models (LLMs) exhibit a wide range of promising capabilities – from step-by-step planning to commonsense reasoning –that provide utility for robot navigation. However, as humans communicate with robots in the real world, ambiguity and uncertainty may be embedded inside spoken instruct…

Cited by 4SourceScholar
2023

AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning

ICRA 2023poster

We propose a novel approach for aerial video action recognition. Our method is designed for videos captured using UAVs and can run on edge or mobile devices. We present a learning-based approach that uses customized auto zoom to automatically identify the human target and scale it appropriately. Thi…

Cited by 18SourceScholar
2023

DroNeRF: Real-Time Multi-Agent Drone Pose Optimization for Computing Neural Radiance Fields

IROS 2023poster

We present a novel optimization algorithm called DroNeRF for the autonomous positioning of monocular camera drones around an object for real-time 3D reconstruction using only a few images. Neural Radiance Fields, or NeRF, is a novel view synthesis technique used to generate new views of an object or…

Cited by 3SourceScholar
2023

EWareNet: Emotion-Aware Pedestrian Intent Prediction and Adaptive Spatial Profile Fusion for Social Robot Navigation

ICRA 2023poster

We present EWareNet, a novel intent and affect-aware social robot navigation algorithm among pedestrians. Our approach predicts the trajectory-based pedestrian intent from gait sequence, which is then used for intent-guided navigation taking into account social and proxemic constraints. We propose a…

Cited by 10SourceScholar
2023

RAIST: Learning Risk Aware Traffic Interactions via Spatio-Temporal Graph Convolutional Networks

IROS 2023poster

A key aspect of driving a road vehicle is to interact with other road users, assess their intentions and make riskaware tactical decisions. An intuitive approach to enabling an intelligent automated driving system would be incorporating some aspects of human driving behavior. To this end, we propose…

Cited by 3SourceScholar
2022

3MASSIV: Multilingual, Multimodal and Multi-Aspect Dataset of Social Media Short Videos

CVPR 2022poster

We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from a social media platform. 3MASSIV comprises of 50k short videos (20 seconds average duration) and 100K unlabeled videos in 11 different languages and captures popular sho…

Cited by 11PDFScholar
2022

Learning Unseen Emotions from Gestures via Semantically-Conditioned Zero-Shot Perception with Adversarial Autoencoders

AAAI 2022technical

We present a novel generalized zero-shot algorithm to recognize perceived emotions from gestures. Our task is to map gestures to novel emotion categories not encountered in training. We introduce an adversarial autoencoder-based representation learning that correlates 3D motion-captured gesture sequ…

Cited by 17SourcePDFScholar
2021

Affect2MM: Affective Analysis of Multimedia Content Using Emotion Causality

CVPR 2021poster

We present Affect2MM, a learning method for time-series emotion prediction for multimedia content. Our goal is to automatically capture the varying emotions depicted by characters in real-life human-centric situations and behaviors. We use the ideas from emotion causation theories to computationally…

Cited by 53PDFcodeScholar
2021

Can a Robot Trust You? : A DRL-Based Approach to Trust-Driven Human-Guided Navigation

ICRA 2021poster

Humans are known to construct cognitive maps of their everyday surroundings using a variety of perceptual inputs. As such, when a human is asked for directions to a particular location, their wayfinding capability in converting this cognitive map into directional instructions is challenged. Owing to…

Cited by 17SourceScholar
2020

CMetric: A Driving Behavior Measure using Centrality Functions

IROS 2020poster

We present a new measure, CMetric, to classify driver behaviors using centrality functions. Our formulation combines concepts from computational graph theory and social traffic psychology to quantify and classify the behavior of human drivers. CMetric is used to compute the probability of a vehicle…

Cited by 47SourceScholar
2020

EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's Principle

CVPR 2020poster

We present EmotiCon, a learning-based algorithm for context-aware perceived human emotion recognition from videos and images. Motivated by Frege's Context Principle from psychology, our approach combines three interpretations of context for emotion recognition. Our first interpretation is based on u…

Cited by 177PDFScholar
2020

Forecasting Trajectory and Behavior of Road-Agents Using Spectral Clustering in Graph-LSTMs

RA-L 2020

We present a novel approach for traffic forecasting in urban traffic scenarios using a combination of spectral graph analysis and deep learning. We predict both the low-level information (future trajectories) as well as the high-level information (road-agent behavior) from the extracted trajectory o

Cited by 175SourceScholar
2020

GraphRQI: Classifying Driver Behaviors Using Graph Spectrums

ICRA 2020poster

We present a novel algorithm (GraphRQI) to identify driver behaviors from road-agent trajectories. Our approach assumes that the road-agents exhibit a range of driving traits, such as aggressive or conservative driving. Moreover, these traits affect the trajectories of nearby road-agents as well as…

Cited by 30SourceScholar
2020

ProxEmo: Gait-based Emotion Learning and Multi-view Proxemic Fusion for Socially-Aware Robot Navigation

IROS 2020poster

We present ProxEmo, a novel end-to-end emotion prediction algorithm for socially aware robot navigation among pedestrians. Our approach predicts the perceived emotions of a pedestrian from walking gaits, which is then used for emotion-guided navigation taking into account social and proxemic constra…

Cited by 93SourcecodeScholar
2020

RoadTrack: Realtime Tracking of Road Agents in Dense and Heterogeneous Environments

ICRA 2020poster

We present a realtime tracking algorithm, Road-Track, to track heterogeneous road-agents in dense traffic videos. Our approach is designed for dense traffic scenarios that consist of different road-agents such as pedestrians, two-wheelers, cars, buses, etc. sharing the road. We use the tracking-by-d…

Cited by 11SourceScholar
2020

Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping

ECCV 2020poster

We present an autoencoder-based semi-supervised approach to classify perceived human emotions from walking styles obtained from videos or motion-captured data and represented as sequences of 3D poses. Given the motion on each joint in the pose at each time step extracted from 3D pose sequences, we h…

Cited by 62SourcePDFScholar
2019

DensePeds: Pedestrian Tracking in Dense Crowds Using Front-RVO and Sparse Features

IROS 2019poster

We present a pedestrian tracking algorithm, DensePeds, that tracks individuals in highly dense crowds (>2 pedestrians per square meter). Our approach is designed for videos captured from front-facing or elevated cameras. We present a new motion model called Front-RVO (FRVO) for predicting pedestrian…

Cited by 22SourceScholar
2019

Pedestrian Dominance Modeling for Socially-Aware Robot Navigation

ICRA 2019poster

We present a Pedestrian Dominance Model (PDM) to identify the dominance characteristics of pedestrians for robot navigation. Through a perception study on a simulated dataset of pedestrians, PDM models the perceived dominance levels of pedestrians with varying motion behaviors corresponding to traje…

Cited by 49SourceScholar
2019

TraPHic: Trajectory Prediction in Dense and Heterogeneous Traffic Using Weighted Interactions

CVPR 2019poster

We present a new algorithm for predicting the near-term trajectories of road agents in dense traffic videos. Our approach is designed for heterogeneous traffic, where the road agents may correspond to buses, cars, scooters, bi-cycles, or pedestrians. We model the interactions between different road…

Cited by 347PDFScholar
2018

Identifying Driver Behaviors Using Trajectory Features for Vehicle Navigation

IROS 2018poster

We present a novel approach to automatically identify driver behaviors from vehicle trajectories and use them for safe navigation of autonomous vehicles. We propose a novel set of features that can be easily extracted from car trajectories. We derive a data-driven mapping between these features and…

Cited by 54SourceScholar
2018

PORCA: Modeling and Planning for Autonomous Driving Among Many Pedestrians

RA-L 2018

This letter presents a planning system for autonomous driving among many pedestrians. A key ingredient of our approach is Pedestrian Optimal Reciprocal Collision Avoidance, a pedestrian motion prediction model that accounts for both a pedestrian's global navigation intention and local interactions w

Cited by 191SourceScholar
2018

The Socially Invisible Robot Navigation in the Social World Using Robot Entitativity

IROS 2018poster

We present a real-time, data-driven algorithm to enhance the social-invisibility of robots within crowds. Our approach is based on prior psychological research, which reveals that people notice and-importantly-react negatively to groups of social actors when they have high entitativity, moving in a…

Cited by 24SourceScholar
2017

SocioSense: Robot navigation amongst pedestrians with social and psychological constraints

IROS 2017poster

We present a real-time algorithm, SocioSense, for socially-aware navigation of a robot amongst pedestrians. Our approach computes time-varying behaviors of each pedestrian using Bayesian learning and Personality Trait theory. These psychological characteristics are used for long-term path prediction…

Cited by 117SourceScholar
2016

GLMP- realtime pedestrian path prediction using global and local movement patterns

ICRA 2016

We present a novel real-time algorithm to predict the path of pedestrians in cluttered environments. Our approach makes no assumption about pedestrian motion or crowd density, and is useful for short-term as well as long-term prediction. We interactively learn the characteristics of pedestrian motio

Cited by 75SourceScholar