← Search

Kaushik Roy

51 accepted papers

2026

ASMA: An Adaptive Safety Margin Algorithm for Vision-Language Drone Navigation Via Scene-Aware Control Barrier Functions

ICRA 2026poster

In the rapidly evolving field of vision–language navigation (VLN), ensuring safety for physical agents remains an open challenge. For a human-in-the-loop language-operated drone to navigate safely, it must understand natural language commands, perceive the environment, and simultaneously avoid hazar…

2026

In-Situ Eval: A Modular Framework for Custom and Real-Time RAG Benchmarking

AAAI 2026technical

Retrieval-Augmented Generation (RAG) has become the standard approach for integrating domain knowledge into Large Language Models (LLMs). However, fair comparison of RAG pipelines remains difficult: data preparation is often ad hoc, subsampling methods are opaque, parameters vary across implementati

Cited by 0SourcePDFScholar
2026

Memorization Through the Lens of Sample Gradients

ICLR 2026poster

Deep neural networks are known to often memorize underrepresented, hard examples, with implications for generalization and privacy. Feldman & Zhang (2020) defined a rigorous notion of memorization. However it is prohibitively expensive to compute at scale because it requires training models both w…

Cited by 0SourcecodeScholar
2026

SPREAD: Subspace Representation Distillation for Lifelong Imitation Learning

ICRA 2026poster

A central challenge in lifelong imitation learning (LIL) is enabling agents to acquire new skills from expert demonstrations while retaining knowledge of previously learned tasks. Achieving this requires preserving the low-dimensional manifolds and geometric structures that underlie task representat…

2026

TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning

ICML 2026poster

Instruction tuning is essential for aligning large language models (LLMs) to downstream tasks and commonly relies on large, diverse corpora. However, small, high-quality subsets, known as coresets, can deliver comparable or superior results, though curating them remains challenging. Existing methods…

Cited by 0SourceScholar
2026

WildCross: A Cross-Modal Large Scale Benchmark for Place Recognition and Metric Depth Estimation in Natural Environments

ICRA 2026poster

Recent years have seen a significant increase in demand for robotic solutions in unstructured natural environments, alongside growing interest in bridging 2D and 3D scene understanding. However, existing robotics datasets are predominantly captured in structured urban environments, making them inade…

2025

ASMA: An $\underline{\text{A}}$daptive $\underline{\text{S}}$afety $\underline{\text{M}}$argin $\underline{\text{A}}$lgorithm for Vision-Language Drone Navigation via Scene-Aware Control Barrier Functions

RA-L 2025

In the rapidly evolving field of vision-language navigation (VLN), ensuring safety for physical agents remains an open challenge. For a human-in-the-loop language-operated drone to navigate safely, it must understand natural language commands, perceive the environment, and simultaneously avoid hazar

Cited by 0SourceScholar
2025

CODE-CL: Conceptor-Based Gradient Projection for Deep Continual Learning

ICCV 2025poster

Continual learning (CL) -- the ability to progressively acquire and integrate new concepts -- is essential to intelligent systems to adapt to dynamic environments. However, deep neural networks struggle with catastrophic forgetting (CF) when learning tasks sequentially, as training for new tasks oft…

2025

CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models

NeurIPS 2025spotlight

As Text-to-Image models continue to evolve, so does the risk of generating unsafe, copyrighted, or privacy-violating content. Existing safety interventions - ranging from training data curation and model fine-tuning to inference-time filtering and guidance - often suffer from incomplete concept remo…

Cited by 0SourceScholar
2025

LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models

EMNLP 2025

Large language models (LLMs) have achieved remarkable multi-step reasoning capabilities across various domains. However, LLMs still face distinct challenges in complex logical reasoning, as (1) proof-finding requires systematic exploration and the maintenance of logical coherence and (2) searching t

2025

M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning

ICRA 2025

Lifelong imitation learning for manipulation tasks poses significant challenges due to distribution shifts that occur in incremental learning steps. Existing methods often rely on unsupervised skill discovery to construct an ever-growing skill library or distillation from multiple policies, which ca

Cited by 9SourceScholar
2025

PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery

AAAI 2025technical

Recent advancements in language-guided diffusion models for image editing are often bottle-necked by cumbersome prompt engineering to precisely articulate desired changes. An intuitive alternative calls on guidance from in-the-wild image exemplars to help users bring their imagined edits to life. Co…

2025

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

ICML 2025spotlight

Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due t…

2025

SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise Robustness

AAAI 2025technical

Label corruption, where training samples are mislabeled due to non-expert annotation or adversarial attacks, significantly degrades model performance. Acquiring large, perfectly labeled datasets is costly, and retraining models from scratch is computationally expensive. To address this, we introduce…

2025

Towards Memorization Estimation: Fast, Formal and Free

ICML 2025poster

Deep learning has become the de facto approach in nearly all learning tasks. It has been observed that deep models tend to memorize and sometimes overfit data, which can lead to compromises in performance, privacy, and other critical metrics. In this paper, we explore the theoretical foundations th…

Cited by 0SourcePDFScholar
2024

Best of Both Worlds: Hybrid SNN-ANN Architecture for Event-based Optical Flow Estimation

IROS 2024poster

In the field of robotics, event-based cameras are emerging as a promising low-power alternative to traditional frame-based cameras for capturing high-speed motion and high dynamic range scenes. This is due to their sparse and asynchronous event outputs. Spiking Neural Networks (SNNs) with their asyn…

Cited by 5SourceScholar
2024

Curvature Clues: Decoding Deep Learning Privacy with Input Loss Curvature

NeurIPS 2024spotlight

In this paper, we explore the properties of loss curvature with respect to input data in deep neural networks. Curvature of loss with respect to input (termed input loss curvature) is the trace of the Hessian of the loss with respect to the input. We investigate how input loss curvature varies betwe…

Cited by 2SourcePDFScholar
2024

EV-Planner: Energy-Efficient Robot Navigation via Event-Based Physics-Guided Neuromorphic Planner

RA-L 2024

Vision-based object tracking is an essential precursor to performing autonomous aerial navigation in order to avoid obstacles. Biologically inspired neuromorphic event cameras are emerging as a powerful alternative to frame-based cameras, due to their ability to asynchronously detect varying intensi

Cited by 24SourcecodeScholar
2024

Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

EMNLP 2024finding

Large language models (LLMs) represent a groundbreaking advancement in the domain of natural language processing due to their impressive reasoning abilities. Recently, there has been considerable interest in increasing the context lengths for these models to enhance their applicability to complex ta…

2024

FEDORA: A Flying Event Dataset fOr Reactive behAvior

IROS 2024poster

The ability of resource-constrained biological systems such as fruitflies to perform complex and high-speed maneuvers in cluttered environments has been one of the prime sources of inspiration for developing vision-based autonomous systems. To emulate this capability, the perception pipeline of such…

Cited by 1SourceScholar
2024

GEAR-Up: Generative AI and External Knowledge-Based Retrieval: Upgrading Scholarly Article Searches for Systematic Reviews

AAAI 2024technical

This paper addresses the time-intensive nature of systematic reviews (SRs) and proposes a solution leveraging advancements in Generative AI (e.g., ChatGPT) and external knowledge augmentation (e.g., Retrieval-Augmented Generation). The proposed system, GEAR-Up, automates query development and transl…

Cited by 7SourcePDFScholar
2024

Memorization Through the Lens of Curvature of Loss Function Around Samples

ICML 2024spotlight

Deep neural networks are over-parameterized and easily overfit to and memorize the datasets that they train on. In the extreme case, it has been shown that networks can memorize a randomly labeled dataset. In this paper, we propose using the curvature of the loss function around each training sample…

Cited by 14SourcePDFScholar
2024

Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models

EMNLP 2024finding

Prompt-based learning is susceptible to intrinsic bias present in pre-trained language models (LMs), leading to sub-optimal performance in prompt-based zero/few-shot settings. In this work, we propose a null-input prompting method to calibrate intrinsic bias encoded in pre-trained LMs. Different fro…

Cited by 1SourcePDFScholar
2024

Unveiling Privacy, Memorization, and Input Curvature Links

ICML 2024poster

Deep Neural Nets (DNNs) have become a pervasive tool for solving many emerging problems. However, they tend to overfit to and memorize the training set. Memorization is of keen interest since it is closely related to several concepts such as generalization, noisy learning, and privacy. To study memo…

Cited by 9SourcePDFScholar
2023

Adaptive-SpikeNet: Event-based Optical Flow Estimation using Spiking Neural Networks with Learnable Neuronal Dynamics

ICRA 2023poster

Event-based cameras have recently shown great potential for high-speed motion estimation owing to their ability to capture temporally rich information asynchronously. Spiking Neural Networks (SNNs), with their neuro-inspired event-driven processing can efficiently handle such asynchronous data, whil…

Cited by 32SourceScholar
2023

DOTIE - Detecting Objects through Temporal Isolation of Events using a Spiking Architecture

ICRA 2023poster

Vision-based autonomous navigation systems rely on fast and accurate object detection algorithms to avoid obstacles. Algorithms and sensors designed for such systems need to be computationally efficient, due to the limited energy of the hardware used for deployment. Biologically inspired event camer…

Cited by 27SourcecodeScholar
2023

Demo Alleviate: Demonstrating Artificial Intelligence Enabled Virtual Assistance for Telehealth: The Mental Health Case

AAAI 2023technical

After the pandemic, artificial intelligence (AI) powered support for mental health care has become increasingly important. The breadth and complexity of significant challenges required to provide adequate care involve: (a) Personalized patient understanding, (b) Safety-constrained and medically vali…

Cited by 21SourcePDFScholar
2023

Event-based Temporally Dense Optical Flow Estimation with Sequential Learning

ICCV 2023poster

Event cameras provide an advantage over traditional frame-based cameras when capturing fast-moving objects without a motion blur. They achieve this by recording changes in light intensity (known as events), thus allowing them to operate at a much higher frequency and making them suitable for capturi…

Cited by 17PDFcodeScholar
2023

Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous Data

NeurIPS 2023poster

Decentralized learning enables the training of deep learning models over large distributed datasets generated at different locations, without the need for a central server. However, in practical scenarios, the data distribution across these devices can be significantly different, leading to a degrad…

2023

Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence Model

EMNLP 2023long findings

Transformers have shown dominant performance across a range of domains including language and vision. However, their computational cost grows quadratically with the sequence length, making their usage prohibitive for resource-constrained applications. To counter this, our approach is to divide the w…

Cited by 0SourceScholar
2022

Fusion-FlowNet: Energy-Efficient Optical Flow Estimation using Sensor Fusion and Deep Fused Spiking-Analog Network Architectures

ICRA 2022poster

Standard frame-based cameras that sample light intensity frames are heavily impacted by motion blur for high-speed motion and fail to perceive scene accurately in high-dynamic range environments. Event-based cameras, on the other hand, overcome these limitations by asynchronously detecting the varia…

Cited by 50SourceScholar
2022

Oscillatory Fourier Neural Network: A Compact and Efficient Architecture for Sequential Processing

AAAI 2022technical

Tremendous progress has been made in sequential processing with the recent advances in recurrent neural networks. However, recurrent architectures face the challenge of exploding/vanishing gradients during training, and require significant computational resources to execute back-propagation through…

Cited by 8SourcePDFScholar
2022

RAPID-RL: A Reconfigurable Architecture with Preemptive-Exits for Efficient Deep-Reinforcement Learning

ICRA 2022poster

Present-day Deep Reinforcement Learning (RL) systems show great promise towards building intelligent agents surpassing human-level performance. However, the computational complexity associated with the underlying deep neural networks (DNNs) leads to power-hungry implementations. This makes deep RL s…

Cited by 5SourceScholar
2022

Spiking Neural Networks with Improved Inherent Recurrence Dynamics for Sequential Learning

AAAI 2022technical

Spiking neural networks (SNNs) with leaky integrate and fire (LIF) neurons, can be operated in an event-driven manner and have internal states to retain information over time, providing opportunities for energy-efficient neuromorphic computing, especially on edge devices. Note, however, many represe…

2022

Towards Ultra Low Latency Spiking Neural Networks for Vision and Sequential Tasks Using Temporal Pruning

ECCV 2022poster

"Spiking Neural Networks (SNNs) can be energy efficient alternatives to commonly used deep neural networks (DNNs). However, computation over multiple timesteps increases latency and energy and incurs memory access overhead of membrane potentials. Hence, latency reduction is pivotal to obtain SNNs wi…

Cited by 40SourcePDFScholar
2021

DCT-SNN: Using DCT To Distribute Spatial Information Over Time for Low-Latency Spiking Neural Networks

ICCV 2021poster

Spiking Neural Networks (SNNs) offer a promising alternative to traditional deep learning frameworks, since they provide higher computational efficiency due to event-driven information processing. SNNs distribute the analog values of pixel intensities into binary spikes over time. However, the most…

Cited by 46PDFcodeScholar
2021

Self-Supervised Optical Flow with Spiking Neural Networks and Event Based Cameras

IROS 2021poster

Optical flow can be leveraged in robotic systems for obstacle detection where low latency solutions are critical in highly dynamic settings. While event-based cameras have changed the dominant paradigm of sending by encoding stimuli into spike trails, offering low bandwidth and latency, events are s…

Cited by 19SourceScholar
2020

Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent Backpropagation

ICLR 2020poster

Spiking Neural Networks (SNNs) operate with asynchronous discrete events (or spikes) which can potentially lead to higher energy-efficiency in neuromorphic hardware implementations. Many works have shown that an SNN for inference can be formed by copying the weights from a trained Artificial Neural…

Cited by 393SourcecodeScholar
2020

Inherent Adversarial Robustness of Deep Spiking Neural Networks: Effects of Discrete Input Encoding and Non-Linear Activations

ECCV 2020poster

In the recent quest for trustworthy neural networks, we present Spiking Neural Network (SNN) as a potential candidate for inherent robustness against adversarial attacks. In this work, we demonstrate that adversarial accuracy of SNNs under gradient-based attacks is higher than their non-spiking coun…

2020

RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network

CVPR 2020poster

Spiking Neural Networks (SNNs) have recently attracted significant research interest as the third generation of artificial neural networks that can enable low-power event-driven data analytics. The best performing SNNs for image recognition tasks are obtained by converting a trained Analog Neural Ne…

Cited by 430PDFcodeScholar
2020

Spike-FlowNet: Event-based Optical Flow Estimation with Energy-Efficient Hybrid Neural Networks

ECCV 2020poster

Event-based cameras display great potential for a variety of tasks such as high-speed motion detection and navigation in low-light environments where conventional frame-based cameras suffer critically. This is attributed to their high temporal resolution, high dynamic range, and low-power consumptio…

2020

Training Deep Spiking Neural Networks for Energy-Efficient Neuromorphic Computing

ICASSP 2020accepted

Spiking Neural Networks (SNNs), widely known as the third generation of neural networks, encode input information temporally using sparse spiking events, which can be harnessed to achieve higher computational efficiency for cognitive tasks. However, considering the rapid strides in accuracy enabled…

Cited by 0SourceScholar
2020

Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition

CVPR 2020oral

Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper…

Cited by 62PDFScholar