← Search

Swagat Kumar

9 accepted papers

2026

Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters

CVPR 2026

Recent adapter-based CLIP tuning (e.g., Tip-Adapter) is a strong few-shot learner, achieving efficiency by caching support features for fast prototype matching. However, these methods rely on global uni-modal feature vectors, overlooking fine-grained patch relations and their structural alignment wi

Cited by 0SourcecodeScholar
2022

Attentive One-Shot Meta-Imitation Learning From Visual Demonstration

ICRA 2022poster

The ability to apply a previously-learned skill (e.g., pushing) to a new task (context or object) is an important requirement for new-age robots. An attempt is made to solve this problem in this paper by proposing a deep meta-imitation learning framework comprising of an attentive-embedding net-work…

Cited by 3SourceScholar
2021

Attentional Learn-able Pooling for Human Activity Recognition

ICRA 2021poster

Human activity/behaviour monitoring and recognition is a key for facilitating humans robot interaction, and allows robots for a better scheduling of future operations. It is challenging and often addressed at different levels, such as human activity classification, future activity prediction and mon…

Cited by 7SourceScholar
2020

Attentive Task-Net: Self Supervised Task-Attention Network for Imitation Learning using Video Demonstration

ICRA 2020poster

This paper proposes an end-to-end self-supervised feature representation network named Attentive Task-Net or AT-Net for video-based task imitation. The proposed AT-Net incorporates a novel multi-level spatial attention module to highlight spatial features corresponding to the intended task demonstra…

Cited by 11SourceScholar
2020

Unsupervised Depth and Confidence Prediction from Monocular Images using Bayesian Inference

IROS 2020poster

In this paper, we propose an unsupervised deep learning framework with Bayesian inference for improving the accuracy of per-pixel depth prediction from monocular RGB images. The proposed framework predicts confidence map along with depth and pose information for a given input image. The depth hypoth…

Cited by 7SourceScholar
2020

Unsupervised Monocular Depth Estimation for Night-time Images using Adversarial Domain Feature Adaptation

ECCV 2020poster

In this paper, we look into the problem of estimating per-pixel depth maps from unconstrained RGB monocular night-time images which is a difficult task that has not been addressed adequately in the literature. The state-of-the-art day-time depth estimation methods fail miserably when tested with nig…

2019

Look No Deeper: Recognizing Places from Opposing Viewpoints under Varying Scene Appearance using Single-View Depth Estimation

ICRA 2019poster

Visual place recognition (VPR) - the act of recognizing a familiar visual place - becomes difficult when there is extreme environmental appearance change or viewpoint change. Particularly challenging is the scenario where both phenomena occur simultaneously, such as when returning for the first time…

Cited by 29SourcecodeScholar
2018

UnDEMoN: Unsupervised Deep Network for Depth and Ego-Motion Estimation

IROS 2018poster

This paper presents a deep network based unsupervised visual odometry system for 6-DoF camera pose estimation and finding dense depth map for its monocular view. The proposed network is trained using unlabeled binocular stereo image pairs and is shown to provide superior performance in depth and ego…

Cited by 44SourceScholar
2017

Improving condition- and environment-invariant place recognition with semantic place categorization

IROS 2017poster

The place recognition problem comprises two distinct subproblems; recognizing a specific location in the world (“specific” or “ordinary” place recognition) and recognizing the type of place (place categorization). Both are important competencies for mobile robots and have each received significant a…

Cited by 38SourceScholar