← Search

Dinesh Manocha

244 accepted papers

2026

AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

CVPR 2026

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning, tracking who speaks, maintaining roles, and grounding events across time. These scenarios are central to multim

Cited by 0SourceScholar
2026

EgoAVU: Egocentric Audio-Visual Understanding

CVPR 2026

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent joint-modality information, whether MLLMs can jointly understan

Cited by 0SourcecodeScholar
2026

Exploring Audio Hallucination in Egocentric Video Understanding

ICASSP 2026oral

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement. State-of-the-art large audio-visual language models (AV-LLMs) can gene…

Cited by 0SourcePDFScholar
2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2026

Music Flamingo: Scaling Music Understanding in Audio Language Models

ICLR 2026poster

We introduce Music Flamingo, a novel large audio–language model, designed to advance music (including song) understanding in foundational audio models. While audio–language research has progressed rapidly, music remains challenging due to its dynamic, layered, and information-dense nature. Progress…

Cited by 0SourcecodeScholar
2026

NavMoE: Hybrid Model and Learning-Based Traversability Estimation for Local Navigation Via Mixture of Experts

ICRA 2026poster

This paper explores traversability estimation for robot navigation. A key bottleneck in traversability estimation lies in efficiently achieving reliable and robust predictions while accurately encoding both geometric and semantic information across diverse environments. We introduce Navigation via M…

2026

PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation

CVPR 2026

Understanding physical properties such as friction, stiffness, hardness, and material composition is essential for enabling robots to interact safely and effectively with their surroundings. However, existing 3D reconstruction methods focus on geometry and appearance and cannot infer these underlyin

Cited by 0SourceScholar
2026

Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away

ICML 2026poster

Reinforcement learning (RL) based post-training for explicit chain-of-thought (e.g., GRPO) improves the reasoning ability of multimodal large-scale reasoning models (MLRMs). But recent evidence shows that it can simultaneously degrade safety alignment and increase jailbreak success rates. We propose…

Cited by 0SourceScholar
2026

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian Splatting

AAAI 2026technical

Despite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple small, moving humans, which are not adequately represented in

Cited by 0SourcePDFScholar
2025

AURELIA: Test-time Reasoning Distillation in Audio-Visual LLMs

ICCV 2025poster

Recent advancements in reasoning optimization have greatly enhanced the performance of large language models (LLMs). However, existing work fails to address the complexities of audio-visual scenarios, underscoring the need for further research. In this paper, we introduce AURELIA, a novel actor-crit…

2025

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

ICCV 2025poster

With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multimodal reasoning proficiency. However, these benchmarks are restricted to assessing primarily the visual aspect and do not examine the holis…

Cited by 0SourcePDFScholar
2025

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

ICML 2025poster

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Language Model (ALM) with advanced audio understanding and reasoning capabilities. AF2…

Cited by 9SourcePDFScholar
2025

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

NeurIPS 2025spotlight

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning acros…

Cited by 0SourcecodeScholar
2025

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning

IROS 2025

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs’ spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs auto-labeling, our approach tackles the challenge of VLMs’ limited spatia

Cited by 9SourceScholar
2025

Behav: Behavioral Rule Guided Autonomy Using VLMs for Robot Navigation in Outdoor Scenes

ICRA 2025

We present BehAV, a novel approach for autonomous robot navigation in outdoor scenes guided by human instructions and leveraging Vision Language Models (VLMs). Our method interprets human commands using a Large Language Model (LLM), and categorizes the instructions into navigation and behavioral gui

Cited by 22SourceScholar
2025

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time

ICML 2025poster

Aligning large language models with humans is challenging due to the inherently multifaceted nature of preference feedback. While existing approaches typically frame this as a multi-objective optimization problem, they often overlook how humans actually make decisions. Research on bounded rationalit…

Cited by 0SourcePDFScholar
2025

CROSS-GAiT: Cross-Attention-Based Multimodal Representation Fusion for Parametric Gait Adaptation in Complex Terrains

IROS 2025

We present CROSS-GAiT, a novel algorithm for quadruped robots that uses Cross Attention to fuse terrain representations derived from visual and time-series inputs; including linear accelerations, angular velocities, and joint efforts. These fused representations are used to continuously adjust two c

Cited by 8SourceScholar
2025

CSCPR: Cross-Source-Context Indoor RGB-D Place Recognition

RA-L 2025

We extend our previous work, PoCo (Liang et al. 2024), and present a new algorithm, Cross-Source-Context Place Recognition (CSCPR), for RGB-D indoor place recognition that integrates global retrieval and reranking into an end-to-end model and keeps the consistency of using Context-of-Clusters (CoCs)

Cited by 1SourceScholar
2025

ChartLens: Fine-grained Visual Attribution in Charts

ACL 2025long

The growing capabilities of multimodal large language models (MLLMs) have advanced tasks like chart understanding. However, these models often suffer from hallucinations, where generated text sequences conflict with the provided visual data. To address this, we introduce Post-Hoc Visual Attribution…

Cited by 0SourcePDFScholar
2025

Collab: Controlled Decoding using Mixture of Agents for LLM Alignment

ICLR 2025poster

Alignment of Large Language models (LLMs) is crucial for safe and trustworthy deployment in applications. Reinforcement learning from human feedback (RLHF) has emerged as an effective technique to align LLMs to human preferences, and broader utilities, but it requires updating billions of model para…

Cited by 1SourcePDFScholar
2025

Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation

IROS 2025

Reinforcement learning (RL) is a promising approach for robotic navigation, allowing robots to learn through trial and error. However, real-world robotic tasks often suffer from sparse rewards, leading to inefficient exploration and suboptimal policies due to sample inefficiency of RL. In this work,

Cited by 4SourceScholar
2025

Do Audio-Language Models Understand Linguistic Variations?

NAACL 2025short

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existi…

2025

Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models

NeurIPS 2025poster

Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek R1) have led to a popular belief that extending thinking traces using prompts like “Wait” or “Let me rethink” can improve performance. This raises a natural question: Does thinking more at test-time truly lead to bet…

Cited by 0SourceScholar
2025

EDM: Equirectangular Projection-Oriented Dense Kernelized Feature Matching

CVPR 2025poster

We introduce the first learning-based dense matching algorithm, termed Equirectangular Projection-Oriented Dense Kernelized Feature Matching (EDM), specifically designed for omnidirectional images. Equirectangular projection (ERP) images, with their large fields of view, are particularly suited for…

2025

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EGOILL

2025

ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera

IROS 2025

We introduce ET-Former, a novel end-to-end algorithm for semantic scene completion using a single monocular camera. Our approach generates a semantic occupancy map from single RGB observation while simultaneously providing uncertainty estimates for semantic predictions. By designing a triplane-based

Cited by 3SourcecodeScholar
2025

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

ICCV 2025poster

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this pa…

Cited by 0SourcePDFScholar
2025

Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation

ACL 2025finding

Generative Error Correction (GEC) has emerged as a powerful post-processing method to boost the performance of Automatic Speech Recognition (ASR) systems. In this paper, we first show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting the…

2025

Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents

EMNLP 2025

Flowcharts are a critical tool for visualizing decision-making processes. However, their non-linear structure and complex visual-textual relationships make it challenging to interpret them using LLMs, as vision-language models frequently hallucinate nonexistent connections and decision paths when an

2025

Gnd: Global Navigation Dataset With Multi-Modal Perception and Multi-Category Traversability in Outdoor Campus Environments

ICRA 2025

Navigating large-scale outdoor environments requires complex reasoning in terms of geometric structures, environmental semantics, and terrain characteristics, which are typically captured by onboard sensors such as LiDAR and cameras. While current mobile robots can navigate such environments using p

Cited by 13SourceScholar
2025

HALO : Human Preference Aligned Offline Reward Learning for Robot Navigation

CoRL 2025poster

In this paper, we introduce HALO, a novel Offline Reward Learning algorithm that quantifies human intuition in navigation into a vision-based reward function for robot navigation. HALO learns a reward model from offline data, leveraging expert trajectories collected from mobile robots. During traini…

Cited by 0SourceScholar
2025

HomeEmergency - Using Audio to Find and Respond to Emergencies in the Home

RA-L 2025

In the United States alone accidental home deaths exceed 128,000 per year. Our work aims to enable home robots who respond to emergency scenarios in the home, preventing injuries and deaths. We introduce a new dataset of household emergencies based in the ThreeDWorld simulator. Each scenario in our

Cited by 0SourceScholar
2025

How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings

ICLR 2025poster

Neural networks that map between low dimensional spaces are ubiquitous in computer graphics and scientific computing; however, in their naive implementation, they are unable to learn high frequency information. We present a comprehensive analysis comparing the two most common techniques for mitigati…

Cited by 0SourcePDFScholar
2025

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

CVPR 2025poster

With the widespread deployment of Multimodal Large Language Models (MLLMs) for visual-reasoning tasks, improving their safety has become crucial. Recent research indicates that despite training-time safety alignment, these models remain vulnerable to jailbreak attacks--carefully crafted image-prompt…

Cited by 3SourcePDFScholar
2025

Imposter: Text and Frequency Guidance for Subject Driven Action Personalization using Diffusion Models

COLING 2025main

We present ImPoster, a novel algorithm for generating a target image of a ‘source’ subject performing a ‘driving’ action. The inputs to our algorithm are a single pair of a source image with the subject that we wish to edit and a driving image with a subject of an arbitrary class performing the driv…

2025

Improving Zero-Shot ObjectNav with Generative Communication

ICRA 2025

We propose a new method for improving zero-shot ObjectNav that aims to utilize potentially available environmental percepts for navigational assistance. Our approach takes into account that the ground agent may have limited and sometimes obstructed view. Our formulation encourages Generative Communi

Cited by 1SourceScholar
2025

Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering

IROS 2025

We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and properties ("What is the color of the car?"), situational queries (such as "Is

Cited by 2SourceScholar
2025

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

NeurIPS 2025poster

Large multimodal models (LMMs) have shown remarkable progress in audiovisual understanding, yet they struggle with real-world scenarios that require complex reasoning across extensive video collections. Existing benchmarks for video question answering remain limited in scope, typically involving one…

Cited by 0SourceScholar
2025

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

ICLR 2025spotlight

The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex rea…

Cited by 25SourcePDFScholar
2025

MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

EMNLP 2025

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchm

2025

On the Vulnerability of LLM/VLM-Controlled Robotics

IROS 2025

In this work, we highlight vulnerabilities in robotic systems integrating large language models (LLMs) and vision-language models (VLMs) due to input modality sensitivities. While LLM/VLM-controlled robots show impressive performance across various tasks, their reliability under slight input variati

Cited by 13SourcecodeScholar
2025

PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification

NAACL 2025long

Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. In this paper, we introduce PAT (Parameter-free Audio-Text aligner), a simple and training-free method aimed at boosting zero-shot audio classification performance of CLAP-like ALMs. To achieve t…

2025

ProSE: Diffusion Priors for Speech Enhancement

NAACL 2025long

Speech enhancement (SE) is the fundamental task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning models have been commonly employed for SE, recent research indicates that generative models, such as denoising diffusion…

2025

PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from related Example Banks

NAACL 2025long

Large Language Models (LLMs) have recently demonstrated impressive few-shot learning capabilities through in-context learning (ICL). However, ICL performance is highly dependent on the choice of few-shot demonstrations, making the selection of the most optimal examples a persistent research challeng…

Cited by 0SourcePDFScholar
2025

RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples

EMNLP 2025

Reward models are essential for aligning large language models (LLMs) with human preferences. However, most open-source multilingual reward models are primarily trained on preference datasets in high-resource languages, resulting in unreliable reward signals for low-resource Indic languages. Collect

Cited by 0SourcePDFScholar
2025

RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization

NeurIPS 2025poster

The increasing use of 360$^\circ$ images across various domains has emphasized the need for robust depth estimation techniques tailored for omnidirectional images. However, obtaining large-scale labeled datasets for 360$^\circ$ depth estimation remains a significant challenge. In this paper, we prop…

Cited by 0SourceScholar
2025

ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds

ICASSP 2025accepted

Open-vocabulary audio-language models, like CLAP [1], offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC…

Cited by 0SourceScholar
2025

Social-LLaVA: Enhancing Social Robot Navigation through Human-Language Reasoning

IROS 2025

As mobile robots become increasingly common in human-centric environments, social navigation—adhering to unwritten social norms rather than merely avoiding pedestrians—has drawn growing attention. Existing methods, from hand-crafted techniques to learning-based approaches, often overlook the nuanced

Cited by 5SourceScholar
2025

Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data

ICLR 2025poster

We present Synthio, a novel approach for augmenting small-scale audio classification datasets with synthetic data. Our goal is to improve audio classification accuracy with limited labeled data. Traditional data augmentation techniques, which apply artificial transformations (e.g., adding random noi…

2025

TK-Planes: Tiered K-Planes with High Dimensional Feature Vectors for Dynamic UAV-based Scenes

IROS 2025

In this paper, we present a new approach to improve the neural rendering fidelity of in-the-wild unmanned aerial vehicle (UAV)-based scenes. Our formulation is designed for dynamic scenes, consisting of small moving objects or human actions in particular. We propose an extension of K-Planes Neural R

Cited by 3SourceScholar
2025

Towards Optimal Multi-draft Speculative Decoding

ICLR 2025poster

Large Language Models (LLMs) have become an indispensable part of natural language processing tasks. However, autoregressive sampling has become an efficiency bottleneck. Multi-Draft Speculative Decoding (MDSD) is a recent approach where, when generating each token, a small draft model generates mul…

Cited by 2SourcePDFScholar
2025

VL-TGS: Trajectory Generation and Selection Using Vision Language Models in Mapless Outdoor Environments

RA-L 2025

We present a multi-modal trajectory generation and selection algorithm for real-world mapless outdoor navigation in human-centered environments. Such environments contain rich features like crosswalks, grass, and curbs, which are easily interpretable by humans, but not by mobile robots. We aim to co

Cited by 28SourceScholar
2025

VLM-GroNav: Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments

ICRA 2025

We present a novel autonomous robot navigation algorithm for outdoor environments that is capable of handling diverse terrain traversability conditions. Our approach, VLM-GroNav, uses vision-language models (VLMs) and integrates them with physical grounding that is used to assess intrinsic terrain p

Cited by 15SourceScholar
2025

VLM-Social-Nav: Socially Aware Robot Navigation Through Scoring Using Vision-Language Models

RA-L 2025

We propose VLM-Social-Nav, a novel Vision-Language Model (VLM) based navigation approach to compute a robot's motion in human-centered environments. Our goal is to make real-time decisions on robot actions that are socially compliant with human expectations. We utilize a perception model to detect i

Cited by 75SourceScholar
2025

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

NeurIPS 2025poster

Vision Language models (VLMs) have achieved remarkable success in video understanding tasks. Yet, a key question remains: Do they comprehend visual information or merely learn superficial mappings between visual and textual patterns? Understanding visual cues, particularly those related to physics…

Cited by 0SourcecodeScholar
2025

VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation

NAACL 2025long

Understanding information from a collection of multiple documents, particularly those with visually rich elements, is important for document-grounded question answering. This paper introduces VisDoMBench, the first comprehensive benchmark designed to evaluate QA systems in multi-document settings wi…

2025

Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

ICLR 2025poster

Large Vision-Language Models (LVLMs) often produce responses that misalign with factual information, a phenomenon known as hallucinations. While hallucinations are well-studied, the exact causes behind them remain underexplored. In this paper, we first investigate the root causes of hallucinations i…

2024

"Don't Forget to Put the Milk Back!" Dataset for Enabling Embodied Agents to Detect Anomalous Situations

RA-L 2024

Home robots intend to make their users lives easier. Our work moves toward more helpful home robots by enabling them to inform their users of dangerous or unsanitary anomalies in the home. Some examples of these anomalies include the user leaving their milk out, forgetting to turn off the stove, or

Cited by 13SourceScholar
2024

A Closer Look at the Limitations of Instruction Tuning

ICML 2024poster

Instruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limi…

Cited by 18SourcePDFScholar
2024

ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions

ACL 2024long

We present ABEX, a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks. ABEX is based on ABstract-and-EXpand, a novel paradigm for generating diverse forms of an input document – we first convert a document into its concise, abstra…

2024

AG-Cvg: Coverage Planning with a Mobile Recharging UGV and an Energy-Constrained UAV

ICRA 2024poster

In this paper, we present an approach for coverage path planning for a team of an energy-constrained Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV). Both the UAV and the UGV have predefined areas that they have to cover. The goal is to perform complete coverage by both robots whi…

Cited by 8SourceScholar
2024

AGL-Net: Aerial-Ground Cross-Modal Global Localization with Varying Scales

IROS 2024poster

We present AGL-NET, a novel learning-based method for global localization using LiDAR point clouds and satellite maps. AGL-Net tackles two critical challenges: bridging the representation gap between image and points modalities for robust feature matching, and handling inherent scale discrepancies b…

Cited by 1SourcecodeScholar
2024

AMCO: Adaptive Multimodal Coupling of Vision and Proprioception for Quadruped Robot Navigation in Outdoor Environments

IROS 2024poster

We present AMCO, a novel navigation method for quadruped robots that adaptively combines vision-based and proprioception-based perception capabilities. Our approach uses three cost maps: general knowledge map; traversability history map; and current proprioception map; which are derived from a robot…

Cited by 4SourceScholar
2024

ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations

ACL 2024findings

Neural image classifiers can often learn to make predictions by overly relying on non-predictive features that are spuriously correlated with the class labels in the training data. This leads to poor performance in real-world atypical scenarios where such features are absent. This paper presents ASP…

2024

AV-RIR: Audio-Visual Room Impulse Response Estimation

CVPR 2024poster

Accurate estimation of Room Impulse Response (RIR) which captures an environment's acoustic properties is important for speech processing and AR/VR applications. We propose AV-RIR a novel multi-modal multi-task learning approach to accurately estimate the RIR from a given reverberant speech signal a…

Cited by 14SourcePDFScholar
2024

AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

EMNLP 2024finding

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucina…

2024

Can LLM’s Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis

NAACL 2024short

We present a novel approach to automatically synthesize “wayfinding instructions” for an embodied robot agent. In contrast to prior approaches that are heavily reliant on human-annotated datasets designed exclusively for specific simulation platforms, our algorithm uses in-context learning to condit…

Cited by 7SourcePDFScholar
2024

Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Based Zero-Shot Object Navigation

RA-L 2024

We present language-guided exploration (LGX), a novel algorithm for <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Language-Driven Zero-Shot Object Goal Navigation</i> (L-ZSON), where an embodied agent navigates to an <italic xmlns:mml="http://www.w

Cited by 144SourceScholar
2024

CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP

NAACL 2024findings

We present CoDa (**Co**nstrained Generation based **Da**ta Augmentation), a controllable, effective, and *training-free* data augmentation technique for low-resource (data-scarce) NLP. Our approach is based on prompting off-the-shelf instruction-following Large Language Models (LLMs) for generating…

2024

CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments

IROS 2024

We present CoNVOI, a novel method for autonomous robot navigation in real-world indoor and outdoor environments using Vision Language Models (VLMs). We employ VLMs in two ways: first, we leverage their zero-shot image classification capability to identify the context or scenario (e.g., indoor corrid

Cited by 52SourceScholar
2024

CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

ICLR 2024poster

A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot a…

Cited by 12SourcePDFScholar
2024

DOC-RAG: ASR Language Model Personalization with Domain-Distributed Co-occurrence Retrieval Augmentation

COLING 2024main

We propose DOC-RAG - Domain-distributed Co-occurrence Retrieval Augmentation for ASR language model personalization aiming to improve the automatic speech recognition of rare word patterns in unseen domains. Our approach involves contrastively training a document retrieval module to rank external kn…

Cited by 2SourcePDFScholar
2024

DTG : Diffusion-based Trajectory Generation for Mapless Global Navigation

IROS 2024poster

We present a novel end-to-end diffusion-based trajectory generation method, DTG, for mapless global navigation in challenging outdoor scenarios with occlusions and unstructured off-road features like grass, buildings, bushes, etc. Given a distant goal, our approach computes a trajectory that satisfi…

Cited by 32SourcecodeScholar
2024

Do Vision-Language Models Understand Compound Nouns?

NAACL 2024short

Open-vocabulary vision-language models (VLMs) like CLIP, trained using contrastive loss, have emerged as a promising new paradigm for text-to-image retrieval. However, do VLMs understand compound nouns (CNs) (e.g., *lab coat*) as well as they understand nouns (e.g., *lab*)? We curate Compun, a novel…

2024

DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding

EMNLP 2024main

Document structure editing involves manipulating localized textual, visual, and layout components in document images based on the user’s requests. Past works have shown that multimodal grounding of user requests in the document image and identifying the accurate structural components and their assoc…

Cited by 2SourcePDFScholar
2024

DocScript: Document-level Script Event Prediction

COLING 2024main

We present a novel task of document-level script event prediction, which aims to predict the next event given a candidate list of narrative events in long-form documents. To enable this, we introduce DocSEP, a challenging dataset in two new domains - contractual documents and Wikipedia articles, whe…

Cited by 1SourcePDFScholar
2024

EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning

EMNLP 2024main

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and ada…

2024

FusDom: Combining in-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised Learning

ICASSP 2024accepted

Continued pre-training (CP) offers multiple advantages, like target domain adaptation and the potential to exploit the continuous stream of unlabeled data available online. However, continued pre-training on out-of-domain distributions often leads to catastrophic forgetting of previously acquired kn…

Cited by 0SourceScholar
2024

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

EMNLP 2024main

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abiliti…

2024

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

CVPR 2024poster

We introduce "HallusionBench" a comprehensive benchmark designed for the evaluation of image-context reasoning. This benchmark presents significant challenges to advanced large visual-language models (LVLMs) such as GPT-4V(ision) Gemini Pro Vision Claude 3 and LLaVA-1.5 by emphasizing nuanced unders…

2024

IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning

EMNLP 2024main

Image-text contrastive models such as CLIP learn transferable and robust representations for zero-shot transfer to a variety of downstream tasks. However, to obtain strong downstream performances, prompts need to be carefully curated, which can be a tedious engineering task. To address the issue of…

Cited by 2SourcePDFScholar
2024

LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments

IROS 2024poster

Navigating robots through unstructured terrains is challenging, primarily due to the dynamic environmental changes. While humans adeptly navigate such terrains by using context from their observations, creating a similar context-aware navigation system for robots is difficult. The essence of the iss…

Cited by 11SourceScholar
2024

LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-time Rendering

CVPR 2024poster

Advancements in neural signed distance fields (SDFs) have enabled modeling 3D surface geometry from a set of 2D images of real-world scenes. Baking neural SDFs can extract explicit mesh with appearance baked into texture maps as neural features. The baked meshes still have a large memory footprint a…

Cited by 7SourcePDFScholar
2024

MIM: Indoor and Outdoor Navigation in Complex Environments Using Multi-Layer Intensity Maps

ICRA 2024poster

We present MIM (Multi-Layer Intensity Map), a novel 3D object representation for robot perception and autonomous navigation. MIMs consist of multiple stacked layers of 2D grid maps each derived from reflected point cloud intensities corresponding to a certain height interval. The different layers of…

Cited by 5SourceScholar
2024

MTG: Mapless Trajectory Generator with Traversability Coverage for Outdoor Navigation

ICRA 2024poster

We present a novel learning-based trajectory generation algorithm for outdoor robot navigation. Our goal is to compute collision-free paths that also satisfy the environment-specific traversability constraints. Our approach is designed for global planning using limited onboard robot perception in ma…

Cited by 10SourceScholar
2024

MaxMin-RLHF: Alignment with Diverse Human Preferences

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data. However, the single reward model overlooks the rich diversity of human preferences inherent in data collected from multiple users. In this…

Cited by 82SourcePDFScholar
2024

MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models

CVPR 2024highlight

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media ranging from movies to social media posts. Machine learning models that can synthesize music are predominantly conditioned on textual descriptions of it. Inspi…

2024

Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

ECCV 2024poster

"Leveraging Large Language Models’ remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding o…

2024

PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback

ICLR 2024poster

We present a novel unified bilevel optimization-based framework, \textsf{PARL}, formulated to address the recently highlighted critical issue of policy alignment in reinforcement learning using utility or preference-based feedback. We identify a major gap within current algorithmic designs for solvi…

Cited by 29SourcePDFScholar
2024

PoCo: Point Context Cluster for RGBD Indoor Place Recognition

IROS 2024

We present a novel end-to-end algorithm (PoCo) for the indoor RGB-D place recognition task, aimed at identifying the most likely match for a given query frame within a reference database. The task presents inherent challenges attributed to the constrained field of view and limited range of perceptio

Cited by 2SourcecodeScholar
2024

Position: On the Possibilities of AI-Generated Text Detection

ICML 2024poster

Our study addresses the challenge of distinguishing human-written text from Large Language Model (LLM) outputs. We provide evidence that this differentiation is consistently feasible, except when human and machine text distributions are indistinguishable across their entire support. Employing inform…

Cited by 4SourcePDFScholar
2024

ProNav: Proprioceptive Traversability Estimation for Legged Robot Navigation in Outdoor Environments

RA-L 2024

We propose a novel method, ProNav, which uses proprioceptive signals for traversability estimation in challenging outdoor terrains for autonomous legged robot navigation. Our approach uses sensor data from a legged robot's joint encoders, force, and current sensors to measure the joint positions, fo

Cited by 26SourceScholar
2024

Recap: Retrieval-Augmented Audio Captioning

ICASSP 2024accepted

We present RECAP (REtrieval-Augmented Audio CAPtioning), a novel and effective audio captioning system that generates captions conditioned on an input audio and other captions similar to the audio retrieved from a datastore. Additionally, our proposed method can transfer to any domain without the ne…

Cited by 0SourceScholar
2024

SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition

IROS 2024poster

We present a new learning approach, Soft Conditional Prompt Learning (SCP), which leverages the strengths of prompt learning for aerial video action recognition. Our approach is designed to predict the action of each agent by helping the models focus on the descriptions or instructions associated wi…

Cited by 1SourceScholar
2024

Saliency-Aware Interpolative Augmentation for Multimodal Financial Prediction

COLING 2024main

Predicting price variations of financial instruments for risk modeling and stock trading is challenging due to the stochastic nature of the stock market. While recent advancements in the Financial AI realm have expanded the scope of data and methods they use, such as textual and audio cues from fina…

2024

Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition

ICASSP 2024accepted

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition (ASR). This paper proposes Stable Distillation, a simple and novel approach for SSL-based continued pre-training that b…

Cited by 0SourceScholar
2024

TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing

ACL 2024findings

Given a source and its edited version performed based on human instructions in natural language, how do we extract the underlying edit operations, to automatically replicate similar edits on other images? This is the problem of reverse designing, and we present TAME-RD, a model to solve this problem…

Cited by 0SourcePDFScholar
2024

Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles

ICML 2024poster

In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient me…

Cited by 2SourcePDFScholar
2024

Transfer Q-star : Principled Decoding for LLM Alignment

NeurIPS 2024poster

Aligning foundation models is essential for their safe and trustworthy deployment. However, traditional fine-tuning methods are computationally intensive and require updating billions of model parameters. A promising alternative, alignment via decoding, adjusts the response distribution directly wit…

Cited by 20SourcePDFScholar
2024

UAV-Sim: NeRF-based Synthetic Data Generation for UAV-based Perception

ICRA 2024poster

Tremendous variations coupled with large degrees of freedom in UAV-based imaging conditions lead to a significant lack of data in adequately learning UAV-based perception models. Using various synthetic renderers in conjunction with perception models is prevalent to create synthetic data to augment…

Cited by 11SourceScholar
2024

Unconstrained Model Predictive Control for Robot Navigation under Uncertainty

ICRA 2024poster

In this paper, we present a probabilistic and unconstrained model predictive control formulation for robot navigation under uncertainty. We present (1) a closed-form approximation of the probability of collision that naturally models the propagation of uncertainty over the planning horizon and is co…

Cited by 2SourceScholar
2024

V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation

ECCV 2024poster

"We introduce V-Trans4Style, an innovative algorithm tailored for dynamic video content editing needs. It is designed to adapt videos to different production styles like documentaries, dramas, feature films, or a specific YouTube channel’s video-making technique. Our algorithm recommends optimal vis…

Cited by 0SourcePDFScholar
2024

VAPOR: Legged Robot Navigation in Unstructured Outdoor Environments using Offline Reinforcement Learning

ICRA 2024poster

We present VAPOR, a novel method for autonomous legged robot navigation in unstructured, densely vegetated outdoor environments using offline Reinforcement Learning (RL). Our method trains a novel RL policy using an actor-critic network and arbitrary data collected in real outdoor vegetation. Our po…

Cited by 5SourceScholar
2024

VLPG-Nav: Object Navigation Using Visual Language Pose Graph and Object Localization Probability Maps

IROS 2024poster

We present VLPG-Nav, a visual language navigation method for guiding robots to specified objects within household scenes. Unlike existing methods primarily focused on navigating the robot toward objects, our approach considers the additional challenge of centering the object within the robot’s camer…

Cited by 1SourceScholar
2024

When, What, and with Whom to Communicate: Enhancing RL-based Multi-Robot Navigation through Selective Communication

IROS 2024poster

Decentralized navigation methods rely primarily on local observations, lacking the global awareness needed to coordinate effectively within a multi-agent system. Exchanging relevant messages between agents can promote cooperation and improve navigation efficiency. We present a Reinforcement Learning…

Cited by 3SourceScholar
2023

3D-Online Generalized Sensed Shape Expansion: A Probabilistically Complete Motion Planner in Obstacle-Cluttered Unknown Environments

RA-L 2023

We present an online motion planning algorithm (3D-OGSSE) for generating smooth, collision-free trajectories over multiple planning iterations for a 3-D agent operating in an unknown, obstacle-cluttered, 3-D environment. In each planning iteration, 3D-OGSSE constructs an obstacle-free region termed

Cited by 3SourceScholar
2023

ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER

ACL 2023long

Complex Named Entity Recognition (NER) is the task of detecting linguistically complex named entities in low-context text. In this paper, we present ACLM Attention-map aware keyword selection for Conditional Language Model fine-tuning), a novel data augmentation approach based on conditional generat…

2023

AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning

ICRA 2023poster

We propose a novel approach for aerial video action recognition. Our method is designed for videos captured using UAVs and can run on edge or mobile devices. We present a learning-based approach that uses customized auto zoom to automatically identify the human target and scale it appropriately. Thi…

Cited by 18SourceScholar
2023

AdVerb: Visually Guided Audio Dereverberation

ICCV 2023poster

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our approach incorporates the complementary visual modality to perform audio dereverber…

Cited by 10PDFScholar
2023

AdaptiveON: Adaptive Outdoor Local Navigation Method for Stable and Reliable Actions

RA-L 2023

We present a novel outdoor navigation algorithm to generate stable and efficient actions to navigate a robot to reach a goal. We use a multi-stage training pipeline and show that our approach produces policies that result in stable and reliable robot navigation on complex terrains. Based on the Prox

Cited by 20SourceScholar
2023

Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic

ICML 2023poster

Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mixes exponentially fast with a rate parameter that appears in the step-size selection. Unfortunately, this assumption is…

Cited by 14SourcePDFScholar
2023

CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network

EMNLP 2023long main

The tremendous growth of social media users interacting in online conversations has led to significant growth in hate speech affecting people from various demographics. Most of the prior works focus on detecting explicit hate speech, which is overt and leverages hateful phrases, with very little wor…

Cited by 0SourcecodeScholar
2023

CrossLoc3D: Aerial-Ground Cross-Source 3D Place Recognition

ICCV 2023poster

We present CrossLoc3D, a novel 3D place recognition method that solves a large-scale point matching problem in a cross-source setting. Cross-source point cloud data corresponds to point sets captured by depth sensors with different accuracies or from different distances and perspectives. We address…

Cited by 7PDFcodeScholar
2023

DALE: Generative Data Augmentation for Low-Resource Legal NLP

EMNLP 2023long main

We present DALE, a novel and effective generative Data Augmentation framework for low-resource LEgal NLP. DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents - legal language, with its specialized vocabulary and complex semantics, morp…

Cited by 0SourcecodeScholar
2023

DS-MPEPC: Safe and Deadlock-Avoiding Robot Navigation in Cluttered Dynamic Scenes

IROS 2023poster

We present an algorithm for safe robot navigation in complex dynamic environments using a variant of model predictive equilibrium point control. We use an optimization formulation to navigate robots gracefully in dynamic environments by optimizing over a trajectory cost function at each timestep. We…

Cited by 6SourceScholar
2023

Dealing with Sparse Rewards in Continuous Control Robotics via Heavy-Tailed Policy Optimization

ICRA 2023poster

In this paper, we present a novel Heavy-Tailed Stochastic Policy Gradient (HT-PSG) algorithm to deal with the challenges of sparse rewards in continuous control problems. Sparse rewards are common in continuous control robotics tasks such as manipulation and navigation and make the learning problem…

Cited by 3SourceScholar
2023

DifFAR: Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition

ICRA 2023poster

We present a learning algorithm, DifFAR, for human activity recognition in videos. Our approach is designed for UAV videos, which are mainly acquired from obliquely placed dynamic cameras that contain a human actor along with background motion. Typically, the human actors occupy less than one-tenth…

Cited by 5SourceScholar
2023

DocEdit: Language-Guided Document Editing

AAAI 2023technical

Professional document editing tools require a certain level of expertise to perform complex edit operations. To make editing tools accessible to increasingly novice users, we investigate intelligent document assistant systems that can make or suggest edits based on a user's natural language request.…

Cited by 5SourcePDFScholar
2023

GrASPE: Graph Based Multimodal Fusion for Robot Navigation in Outdoor Environments

RA-L 2023

We present a novel trajectory traversability estimation and planning algorithm for robot navigation in complex outdoor environments. We incorporate multimodal sensory inputs from an RGB camera, 3D LiDAR, and the robot's odometry sensor to train a prediction model to estimate candidate trajectories'

Cited by 78SourceScholar
2023

Intent-Aware Planning in Heterogeneous Traffic via Distributed Multi-Agent Reinforcement Learning

CoRL 2023oral

Navigating safely and efficiently in dense and heterogeneous traffic scenarios is challenging for autonomous vehicles (AVs) due to their inability to infer the behaviors or intentions of nearby drivers. In this work, we introduce a distributed multi-agent reinforcement learning (MARL) algorithm for…

Cited by 10SourceScholar
2023

LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference

ICCV 2023poster

We propose a novel method, LoLep, which regresses Locally-Learned planes from a single RGB image to represent scenes accurately, thus generating better novel views. Without the depth information, regressing appropriate plane locations is a challenging problem. To solve this issue, we pre-partition t…

Cited by 0PDFcodeScholar
2023

METEOR: A Dense, Heterogeneous, and Unstructured Traffic Dataset with Rare Behaviors

ICRA 2023poster

We present a new traffic dataset, Meteor, which captures traffic patterns and multi-agent driving behaviors in unstructured scenarios. Meteor consists of more than 1000 one-minute videos, over 2 million annotated frames with bounding boxes and GPS trajectories for 16 unique agent categories, and mor…

Cited by 15SourceScholar
2023

PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval Augmentation

EMNLP 2023long findings

We introduce PersonaLM - Domain-distributed Span-Aggregated K-nearest N-gram retrieval augmentation to improve language modeling for Automatic Speech Recognition (ASR) personalization. PersonaLM leverages contextually similar n-gram word frequencies for recognizing rare word patterns associated with…

Cited by 0SourceScholar
2023

Posterior Coreset Construction with Kernelized Stein Discrepancy for Model-Based Reinforcement Learning

AAAI 2023technical

Model-based approaches to reinforcement learning (MBRL) exhibit favorable performance in practice, but their theoretical guarantees in large spaces are mostly restricted to the setting when transition model is Gaussian or Lipschitz, and demands a posterior estimate whose representational complexity…

Cited by 11SourcePDFScholar
2023

RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments

ICRA 2023poster

We present a novel reinforcement learning based algorithm for multi-robot task allocation problem in ware-house environments. We formulate it as a Markov Decision Process and solve via a novel deep multi-agent reinforcement learning method (called RTAW) with attention inspired policy architecture. H…

Cited by 29SourceScholar
2023

Real-Time Decentralized Navigation of Nonholonomic Agents Using Shifted Yielding Areas

ICRA 2023poster

We present a lightweight, decentralized algorithm for navigating multiple nonholonomic agents through challenging environments with narrow passages. Our key idea is to allow agents to yield to each other in large open areas instead of narrow passages, to increase the success rate of conventional dec…

Cited by 2SourceScholar
2023

SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training

ICASSP 2023accepted

We present a new Self-Supervised Learning (SSL) approach to pre-train encoders on unlabeled audio data that reduces the need for large amounts of labeled data for audio and speech classification. Our primary aim is to learn au-dio representations that can generalize across a large vari-ety of speech…

Cited by 0SourceScholar
2023

STEERING : Stein Information Directed Exploration for Model-Based Reinforcement Learning

ICML 2023poster

Directed Exploration is a crucial challenge in reinforcement learning (RL), especially when rewards are sparse. Information-directed sampling (IDS), which optimizes the information ratio, seeks to do so by augmenting regret with information gain. However, estimating information gain is computational…

Cited by 8SourcePDFScholar
2023

Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline Performances

ICRA 2023poster

Human action recognition is a challenging problem, particularly when there is high variability in factors such as subject appearance, backgrounds and viewpoint. While deep neural networks (DNNs) have been shown to perform well on action recognition tasks, they typically require large amounts of high…

Cited by 36SourcecodeScholar
2023

TMO: Textured Mesh Acquisition of Objects With a Mobile Device by Using Differentiable Rendering

CVPR 2023poster

We present a new pipeline for acquiring a textured mesh in the wild with a single smartphone which offers access to images, depth maps, and valid poses. Our method first introduces an RGBD-aided structure from motion, which can yield filtered depth maps and refines camera poses guided by correspondi…

Cited by 11SourcePDFScholar
2023

Towards Improved Room Impulse Response Estimation for Speech Recognition

ICASSP 2023accepted

We propose a novel approach for blind room impulse response (RIR) estimation systems in the context of a downstream application scenario, far-field automatic speech recognition (ASR). We first draw the connection between improved RIR estimation and improved ASR performance, as a means of evaluating…

Cited by 0SourceScholar
2023

VERN: Vegetation-Aware Robot Navigation in Dense Unstructured Outdoor Environments

IROS 2023poster

We propose a novel method for autonomous legged robot navigation in densely vegetated environments with a variety of pliable/traversable and non-pliable/untraversable vegetation. We present a novel few-shot learning classifier that can be trained on a few hundred RGB images to differentiate flora th…

Cited by 18SourceScholar
2022

3MASSIV: Multilingual, Multimodal and Multi-Aspect Dataset of Social Media Short Videos

CVPR 2022poster

We present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from a social media platform. 3MASSIV comprises of 50k short videos (20 seconds average duration) and 100K unlabeled videos in 11 different languages and captures popular sho…

Cited by 11PDFScholar
2022

A Repulsive Force Unit for Garment Collision Handling in Neural Networks

ECCV 2022poster

"Despite recent success, deep learning-based methods for predicting 3D garment deformation under body motion suffer from interpenetration problems between the garment and the body. To address this problem, we propose a novel collision handling neural network layer called Repulsive Force Unit (ReFU).…

Cited by 15SourcePDFScholar
2022

AFR: An Efficient Buffering Algorithm for Cloud Robotic Systems

IROS 2022poster

Communication between robots and the server is a major problem for cloud robotic systems. In this paper, we address the problem caused by data loss during such communications and propose an efficient buffering algorithm, called AFR, to solve the problem. We model the problem into an optimization pro…

Cited by 0SourcecodeScholar
2022

B-GAP: Behavior-Rich Simulation and Navigation for Autonomous Driving

RA-L 2022

We address the problem of ego-vehicle navigation in dense simulated traffic environments populated by road agents with varying driver behaviors. Navigation in such environments is challenging due to unpredictability in agents&#x2019; actions caused by their heterogeneous behaviors. We present a new

Cited by 32SourcecodeScholar
2022

CGLR: Dense Multi-Agent Navigation Using Voronoi Cells and Congestion Metric-based Replanning

IROS 2022poster

We present a decentralized path-planning algorithm for navigating multiple differential-drive robots in dense environments. In contrast to prior decentralized methods, we propose a novel congestion metric-based replanning that couples local and global planning techniques to efficiently navigate in s…

Cited by 2SourceScholar
2022

CoMet: Modeling Group Cohesion for Socially Compliant Robot Navigation in Crowded Scenes

RA-L 2022

We present CoMet, a novel approach for computing a group’s cohesion and using that to improve a robot’s navigation in crowded scenes. Our approach uses a novel cohesion-metric that builds on prior work in social psychology. We compute this metric by utilizing various visual features of pedestrians f

Cited by 24SourceScholar
2022

D2-TPred: Discontinuous Dependency for Trajectory Prediction under Traffic Lights

ECCV 2022poster

"A profound understanding of inter-agent relationships and motion behaviors is important to achieve high-quality planning when navigating in complex scenarios, especially at urban traffic intersections. We present a trajectory prediction approach with respect to traffic lights, D2-TPred, which uses…

2022

DC-MRTA: Decentralized Multi-Robot Task Allocation and Navigation in Complex Environments

IROS 2022poster

We present a novel reinforcement learning (RL) based task allocation and decentralized navigation algorithm for mobile robots in warehouse environments. Our approach is designed for scenarios in which multiple robots are used to perform various pick up and delivery tasks. We consider the problem of…

Cited by 17SourceScholar
2022

DocFin: Multimodal Financial Prediction and Bias Mitigation using Semi-structured Documents

EMNLP 2022finding

Financial prediction is complex due to the stochastic nature of the stock market. Semi-structured financial documents present comprehensive financial data in tabular formats, such as earnings, profit-loss statements, and balance sheets, and can often contain rich technical analysis along with a text…

2022

DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection

EMNLP 2022main

We present DocInfer - a novel, end-to-end Document-level Natural Language Inference model that builds a hierarchical document graph enriched through inter-sentence relations (topical, entity-based, concept-based), performs paragraph pruning using the novel SubGraph Pooling layer, followed by optimal…

2022

DocTime: A Document-level Temporal Dependency Graph Parser

NAACL 2022long

We introduce DocTime - a novel temporal dependency graph (TDG) parser that takes as input a text document and produces a temporal dependency graph. It outperforms previous BERT-based solutions by a relative 4-8% on three datasets from modeling the problem as a graph network with path-prediction loss…

2022

FAR: Fourier Aerial Video Recognition

ECCV 2022poster

"We present a method, Fourier Activity Recognition (FAR), for UAV video activity recognition. Our formulation uses a novel Fourier object disentanglement method to innately separate out the human agent (which is typically small) from the background. Our disentanglement technique operates in the freq…

2022

Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator

ICASSP 2022accepted

We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker positions, and reverberation time (T <inf xmlns:mml="http://ww…

Cited by 0SourceScholar
2022

GA-Nav: Efficient Terrain Segmentation for Robot Navigation in Unstructured Outdoor Environments

RA-L 2022

We propose GA-Nav, a novel group-wise attention mechanism to identify safe and navigable regions in unstructured environments from RGB images. Our group-wise attention method extracts multi-scale features from each type of terrain independently and classifies terrains based on their navigability lev

Cited by 161SourcecodeScholar
2022

Game-Theoretic Planning for Autonomous Driving among Risk-Aware Human Drivers

ICRA 2022poster

We present a novel approach for risk-aware planning with human agents in multi-agent traffic scenarios. Our approach takes into account the wide range of human driver behaviors on the road, from aggressive maneuvers like speeding and overtaking, to conservative traits like driving slowly and conform…

Cited by 14SourceScholar
2022

GamePlan: Game-Theoretic Multi-Agent Planning With Human Drivers at Intersections, Roundabouts, and Merging

RA-L 2022

We present a new method for multi-agent planning involving human drivers and autonomous vehicles (AVs) in unsignaled intersections, roundabouts, and during merging. In multi-agent planning, the main challenge is to predict the actions of other agents, especially human drivers, as their intentions ar

Cited by 62SourceScholar
2022

HTRON: Efficient Outdoor Navigation with Sparse Rewards via Heavy Tailed Adaptive Reinforce Algorithm

CoRL 2022poster

We present a novel approach to improve the performance of deep reinforcement learning (DRL) based outdoor robot navigation systems. Most, existing DRL methods are based on carefully designed dense reward functions that learn the efficient behavior in an environment.  We circumvent this issue by work…

Cited by 13SourceScholar
2022

Image-Goal Navigation in Complex Environments via Modular Learning

RA-L 2022

We present a novel approach for image-goal navigation, where an agent navigates with a goal image rather than accurate target information, which is more challenging. Our goal is to decouple the learning of navigation goal planning, collision avoidance, and navigation ending prediction, which enables

Cited by 19SourceScholar
2022

MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints

ICRA 2022poster

We present a novel self-supervised algorithm named MotionHint for monocular visual odometry (VO) that takes motion constraints into account. A key aspect of our approach is to use an appropriate motion model that can help existing self-supervised monocular VO (SSM-VO) algorithms to overcome issues r…

Cited by 13SourcecodeScholar
2022

Multi-Robot Path Planning Using Medial-Axis-Based Pebble-Graph Embedding

IROS 2022poster

We present a centralized algorithm for labeled, disk-shaped Multi-Robot Path Planning (MPP) in a continuous planar workspace with polygonal boundaries. Our method automatically transform the continuous problem into a discrete, graph-based variant termed the pebble motion problem, which can be solved…

Cited by 3SourceScholar
2022

N-Penetrate: Active Learning of Neural Collision Handler for Complex 3D Mesh Deformations

ICML 2022spotlight

We present a robust learning algorithm to detect and handle collisions in 3D deforming meshes. We first train a neural network to detect collisions and then use a numerical optimization algorithm to resolve penetrations guided by the network. Our learned collision handler can resolve collisions for…

Cited by 4SourcePDFScholar
2022

STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes

CVPR 2022poster

Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations w…

Cited by 49PDFcodeScholar
2022

SelfTune: Metrically Scaled Monocular Depth Estimation through Self-Supervised Learning

ICRA 2022poster

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with proprioceptive sensors. Such monocular SLAM systems can provide metric…

Cited by 5SourceScholar
2022

TERP: Reliable Planning in Uneven Outdoor Environments using Deep Reinforcement Learning

ICRA 2022poster

We present a novel method for reliable robot navigation in uneven outdoor terrains. Our approach employs a fully-trained Deep Reinforcement Learning (DRL) network that uses elevation maps of the environment, robot pose, and goal as inputs to compute an attention mask of the environment. The attentio…

Cited by 82SourceScholar
2022

TNS: Terrain Traversability Mapping and Navigation System for Autonomous Excavators

RSS 2022poster

We present a terrain traversability mapping and navigation system (TNS) for autonomous excavator applications in an unstructured environment. We use an efficient approach to extract terrain features from RGB images and 3D point clouds and incorporate them into a global map for planning and navigatio…

2022

TerraPN: Unstructured Terrain Navigation using Online Self-Supervised Learning

IROS 2022poster

We present TerraPN, a novel method that learns the surface properties (traction, bumpiness, deformability, etc.) of complex outdoor terrains directly from robot-terrain interactions through self-supervised learning, and uses it for autonomous robot navigation. Our method uses RGB images of terrain s…

Cited by 62SourceScholar
2021

Affect2MM: Affective Analysis of Multimedia Content Using Emotion Causality

CVPR 2021poster

We present Affect2MM, a learning method for time-series emotion prediction for multimedia content. Our goal is to automatically capture the varying emotions depicted by characters in real-life human-centric situations and behaviors. We use the ideas from emotion causation theories to computationally…

Cited by 53PDFcodeScholar
2021

DWA-RL: Dynamically Feasible Deep Reinforcement Learning Policy for Robot Navigation among Mobile Obstacles

ICRA 2021poster

We present a novel Deep Reinforcement Learning (DRL) based policy to compute dynamically feasible and spatially aware velocities for a robot navigating among mobile obstacles. Our approach combines the benefits of the Dynamic Window Approach (DWA) in terms of satisfying the robot’s dynamics constrai…

Cited by 88SourceScholar
2021

DnD: Dense Depth Estimation in Crowded Dynamic Indoor Scenes

ICCV 2021poster

We present a novel approach for estimating depth from a monocular camera as it moves through complex and crowded indoor environments, e.g., a department store or a metro station. Our approach predicts absolute scale depth maps over the entire scene consisting of a static background and multiple movi…

Cited by 6PDFScholar
2021

Dynamic Graph Modeling Of Simultaneous EEG And Eye-Tracking Data For Reading Task Identification

ICASSP 2021accepted

We present a new approach, that we call AdaGTCN, for identifying human reader intent from Electroencephalogram (EEG) and Eye movement (EM) data in order to help differentiate between normal reading and task-oriented reading. Understanding the physiological aspects of the reading process (the cogniti…

Cited by 0SourceScholar
2021

HighlightMe: Detecting Highlights From Human-Centric Videos

ICCV 2021poster

We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as poses and faces. We use an autoencoder network equipped wit…

Cited by 10PDFScholar
2021

LCollision: Fast Generation of Collision-Free Human Poses using Learned Non-Penetration Constraints

AAAI 2021technical

We present LCollision, a learning-based method that synthesizes collision-free 3D human poses. At the crux of our approach is a novel deep architecture that simultaneously decodes new human poses from the latent space and predicts colliding body parts. These two components of our architecture are us…

Cited by 14SourcePDFScholar
2021

Multi-Agent Ergodic Coverage in Urban Environments

ICRA 2021poster

An important aspect of dynamic urban coverage is how building collision avoidance is incorporated into the overall coverage mission. We consider a multi-agent urban dynamic coverage problem in which a team of flying agents uses downward facing cameras to observe the street-level environment outside…

Cited by 9SourceScholar
2021

ORBBuf: A Robust Buffering Method for Remote Visual SLAM

IROS 2021poster

The data loss caused by unreliable network seriously impacts the results of remote visual SLAM systems. From our experiment, a loss of less than 1 second of data can cause a visual SLAM algorithm to lose tracking. We present a novel buffering method, ORBBuf, to reduce the impact of data loss on remo…

Cited by 9SourceScholar
2021

Point-based Acoustic Scattering for Interactive Sound Propagation via Surface Encoding

IJCAI 2021poster

We present a novel geometric deep learning method to compute the acoustic scattering properties of geometric objects. Our learning algorithm uses a point cloud representation of objects to compute the scattering properties and integrates them with ray tracing for interactive sound propagation in dyn…

2021

Reinforcement Learning-Based Visual Navigation With Information-Theoretic Regularization

RA-L 2021

To enhance the cross-target and cross-scene generalization of target-driven visual navigation based on deep reinforcement learning (RL), we introduce an information-theoretic regularization term into the RL objective. The regularization maximizes the mutual information between navigation actions and

Cited by 35SourcecodeScholar
2021

Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVIS

ICCV 2021poster

We present a novel approach to robustly detect and perceive vehicles in different camera views as part of a cooperative vehicle-infrastructure system (CVIS). Our formulation is designed for arbitrary camera views and makes no assumptions about intrinsic or extrinsic parameters. First, to deal with m…

Cited by 3PDFcodeScholar
2021

SelfDeco: Self-Supervised Monocular Depth Completion in Challenging Indoor Environments

ICRA 2021poster

We present a novel algorithm for self-supervised monocular depth completion. Our approach is based on training a neural network that requires only sparse depth measurements and corresponding monocular video sequences without dense depth labels. Our self-supervised algorithm is designed for challengi…

Cited by 27SourceScholar
2021

SwarmCCO: Probabilistic Reactive Collision Avoidance for Quadrotor Swarms Under Uncertainty

RA-L 2021

We present decentralized collision avoidance algorithms for quadrotor swarms operating under uncertain state estimation. Our approach exploits the differential flatness property and feedforward linearization to approximate the quadrotor dynamics and performs reciprocal collision avoidance. We accoun

Cited by 14SourceScholar
2021

TIMERS: Document-level Temporal Relation Extraction

ACL 2021short

We present TIMERS - a TIME, Rhetorical and Syntactic-aware model for document-level temporal relation classification in the English language. Our proposed method leverages rhetorical discourse features and temporal arguments from semantic role labels, in addition to traditional local syntactic featu…

2021

Towards Target-Driven Visual Navigation in Indoor Scenes via Generative Imitation Learning

RA-L 2021

We present a target-driven navigation system to improve mapless visual navigation in indoor scenes. Our method takes a multi-view observation of a robot and a target image as inputs at each time step to provide a sequence of actions that move the robot to the target without relying on odometry or GP

Cited by 49SourcecodeScholar
2021

V-RVO: Decentralized Multi-Agent Collision Avoidance using Voronoi Diagrams and Reciprocal Velocity Obstacles

IROS 2021poster

We present a decentralized collision avoidance method for dense environments based on buffered Voronoi cells (BVC) and reciprocal velocity obstacles (RVO). Our approach is designed for scenarios with a large number of agents in close proximity and provides passive-friendly collision avoidance guaran…

Cited by 41SourceScholar
2021

XAI-N: Sensor-based Robot Navigation using Expert Policies and Decision Trees

IROS 2021poster

We present a novel sensor-based learning navigation algorithm to compute a collision-free trajectory for a robot in dense and dynamic environments with moving obstacles or targets. Our approach uses deep reinforcement learning-based expert policy that is trained using a sim2real paradigm. In order t…

Cited by 17SourcecodeScholar
2020

AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points

ECCV 2020poster

Current methods for trajectory prediction operate in supervised manners, and therefore require vast quantities of corresponding ground truth data for training. In this paper, we present a novel, label-free algorithm, AutoTrajectory, for trajectory extraction and prediction to use raw videos directly…

2020

CMetric: A Driving Behavior Measure using Centrality Functions

IROS 2020poster

We present a new measure, CMetric, to classify driver behaviors using centrality functions. Our formulation combines concepts from computational graph theory and social traffic psychology to quantify and classify the behavior of human drivers. CMetric is used to compute the probability of a vehicle…

Cited by 47SourceScholar
2020

Crowd-Steer: Realtime Smooth and Collision-Free Robot Navigation in Densely Crowded Scenarios Trained using High-Fidelity Simulation

IJCAI 2020poster

We present a novel high fidelity 3-D simulator that significantly reduces the sim-to-real gap for collision avoidance in dense crowds using Deep Reinforcement Learning (DRL). Our simulator models realistic crowd and pedestrian behaviors, along with friction, sensor noise and delays in the simulated…

Cited by 0SourcePDFScholar
2020

DCAD: Decentralized Collision Avoidance With Dynamics Constraints for Agile Quadrotor Swarms

RA-L 2020

We present DCAD, a novel, decentralized collision avoidance algorithm for navigating a swarm of quadrotors in dense environments populated with static and dynamic obstacles. Our algorithm relies on the concept of Optimal Reciprocal Collision Avoidance (ORCA) and utilizes a flatness-based Model Predi

Cited by 72SourceScholar
2020

Deep Differentiable Grasp Planner for High-DOF Grippers

RSS 2020poster

We present an end-to-end algorithm for training deep neural networks to grasp novel objects. Our algorithm builds all the essential components of a grasping system using a forward-backward automatic differentiation approach, including the forward kinematics of the gripper, the collision between the…

Cited by 78SourcePDFScholar
2020

DeepMNavigate: Deep Reinforced Multi-Robot Navigation Unifying Local & Global Collision Avoidance

IROS 2020poster

We present a novel algorithm (DeepMNavigate) for global multi-agent navigation in dense scenarios using deep reinforcement learning (DRL). Our approach uses local and global information for each robot from motion information maps. We use a three-layer CNN that takes these maps as input to generate a…

Cited by 28SourceScholar
2020

DenseCAvoid: Real-time Navigation in Dense Crowds using Anticipatory Behaviors

ICRA 2020poster

We present DenseCAvoid, a novel algorithm for navigating a robot through dense crowds and avoiding collisions by anticipating pedestrian behaviors. Our formulation uses visual sensors and a pedestrian trajectory prediction algorithm to track pedestrians in a set of input frames and compute bounding…

Cited by 107SourceScholar
2020

EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's Principle

CVPR 2020poster

We present EmotiCon, a learning-based algorithm for context-aware perceived human emotion recognition from videos and images. Motivated by Frege's Context Principle from psychology, our approach combines three interpretations of context for emotion recognition. Our first interpretation is based on u…

Cited by 177PDFScholar
2020

Forecasting Trajectory and Behavior of Road-Agents Using Spectral Clustering in Graph-LSTMs

RA-L 2020

We present a novel approach for traffic forecasting in urban traffic scenarios using a combination of spectral graph analysis and deep learning. We predict both the low-level information (future trajectories) as well as the high-level information (road-agent behavior) from the extracted trajectory o

Cited by 175SourceScholar
2020

Frozone: Freezing-Free, Pedestrian-Friendly Navigation in Human Crowds

RA-L 2020

We present Frozone, a novel algorithm to deal with the Freezing Robot Problem (FRP) that arises when a robot navigates through dense scenarios and crowds. Our method senses and explicitly predicts the trajectories of pedestrians and constructs a Potential Freezing Zone (PFZ); a spatial zone where th

Cited by 87SourceScholar
2020

GraphRQI: Classifying Driver Behaviors Using Graph Spectrums

ICRA 2020poster

We present a novel algorithm (GraphRQI) to identify driver behaviors from road-agent trajectories. Our approach assumes that the road-agents exhibit a range of driving traits, such as aggressive or conservative driving. Moreover, these traits affect the trajectories of nearby road-agents as well as…

Cited by 30SourceScholar
2020

Improving Reverberant Speech Training Using Diffuse Acoustic Simulation

ICASSP 2020accepted

We present an efficient and realistic geometric acoustic simulation approach for generating and augmenting training data in speech-related machine learning tasks. Our physically-based acoustic simulation method is capable of modeling occlusion, specular and diffuse reflections of sound in complicate…

Cited by 0SourceScholar
2020

Learning Resilient Behaviors for Navigation Under Uncertainty

ICRA 2020poster

Deep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). O…

Cited by 28SourceScholar
2020

Low-Frequency Compensated Synthetic Impulse Responses For Improved Far-Field Speech Recognition

ICASSP 2020accepted

We propose a method for generating low-frequency compensated synthetic impulse responses that improve the performance of farfield speech recognition systems trained on artificially augmented datasets. We design linear-phase filters that adapt the simulated impulse responses to equalization distribut…

Cited by 0SourceScholar