← Search

Gaurav Pandey

17 accepted papers

2026

Insertion Based Sequence Generation with Learnable Order Dynamics

ICML 2026poster

In many domains generating variable length sequences through insertions provides greater flexibility over autoregressive models. However, the action space of insertion models is much larger than that of autoregressive models (ARMs) making the learning challenging. To address this, we incorporate tra…

Cited by 0SourceScholar
2025

Systematic Knowledge Injection into Large Language Models via Diverse Augmentation for Domain-Specific RAG

NAACL 2025findings

Retrieval-Augmented Generation (RAG) has emerged as a prominent method for incorporating domain knowledge into Large Language Models (LLMs). While RAG enhances response relevance by incorporating retrieved domain knowledge in the context, retrieval errors can still lead to hallucinations and incorre…

2024

BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback

ICML 2024poster

Distribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likeliho…

Cited by 3SourcePDFScholar
2024

HealthAlignSumm : Utilizing Alignment for Multimodal Summarization of Code-Mixed Healthcare Dialogues

EMNLP 2024finding

As generative AI progresses, collaboration be-tween doctors and AI scientists is leading to thedevelopment of personalized models to stream-line healthcare tasks and improve productivity.Summarizing doctor-patient dialogues has be-come important, helping doctors understandconversations faster and im…

2024

SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images

IROS 2024poster

This research paper presents an innovative multitask learning framework that allows concurrent depth estimation and semantic segmentation using a single camera. The proposed approach is based on a shared encoder-decoder architecture, which integrates various techniques to improve the accuracy of the…

Cited by 11SourcecodeScholar
2023

Stereo Visual Odometry with Deep Learning-Based Point and Line Feature Matching Using an Attention Graph Neural Network

IROS 2023poster

Robust feature matching forms the backbone for most Visual Simultaneous Localization and Mapping (vSLAM), visual odometry, 3D reconstruction, and Structure from Motion (SfM) algorithms. However, recovering feature matches from texture-poor scenes is a major challenge and still remains an open area o…

Cited by 4SourceScholar
2022

Gaining Insights into Unrecognized User Utterances in Task-Oriented Dialog Systems

EMNLP 2022industry

The rapidly growing market demand for automatic dialogue agents capable of goal-oriented behavior has caused many tech-industry leaders to invest considerable efforts into task-oriented dialog systems. The success of these systems is highly dependent on the accuracy of their intent identification –…

Cited by 7SourcePDFScholar
2022

Localization of a Smart Infrastructure Fisheye Camera in a Prior Map for Autonomous Vehicles

ICRA 2022poster

This work presents a technique for localization of a smart infrastructure node, consisting of a fisheye camera, in a prior map. These cameras can detect objects that are outside the line of sight of the autonomous vehicles (AV) and send that information to AVs using V2X technology. However, in order…

Cited by 6SourceScholar
2022

Mix-and-Match: Scalable Dialog Response Retrieval using Gaussian Mixture Embeddings

EMNLP 2022finding

Embedding-based approaches for dialog response retrieval embed the context-response pairs as points in the embedding space. These approaches are scalable, but fail to account for the complex, many-to-many relationships that exist between context-response pairs. On the other end of the spectrum, ther…

2022

Real-time Full-stack Traffic Scene Perception for Autonomous Driving with Roadside Cameras

ICRA 2022poster

We propose a novel and pragmatic framework for traffic scene perception with roadside cameras. The proposed framework covers a full-stack of roadside perception pipeline for infrastructure-assisted autonomous driving, including object detection, object localization, object tracking, and multi-camera…

Cited by 50SourceScholar
2022

Variational Learning for Unsupervised Knowledge Grounded Dialogs

IJCAI 2022poster

Recent methods for knowledge grounded dialogs generate responses by incorporating information from an external textual document. These methods do not require the exact document to be known during training and rely on the use of a retrieval system to fetch relevant documents from a large index. The d…

2021

Simulated Chats for Building Dialog Systems: Learning to Generate Conversations from Instructions

EMNLP 2021finding

Popular dialog datasets such as MultiWOZ are created by providing crowd workers an instruction, expressed in natural language, that describes the task to be accomplished. Crowd workers play the role of a user and an agent to generate dialogs to accomplish tasks involving booking restaurant tables, c…

2020

Experimental Evaluation of 3D-LIDAR Camera Extrinsic Calibration

IROS 2020poster

In this paper we perform an extensive experimental evaluation of three planar target based 3D-LIDAR camera calibration algorithms, on a sensor suite consisting multiple 3D-LIDARs and cameras, assessing their robustness to random initialization and by using metrics like Mean Line Re-projection Error…

Cited by 21SourceScholar