← Search

Kamal Gupta

20 accepted papers

2024

EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS

ECCV 2024poster

"Recently, 3D Gaussian splatting (3D-GS) has gained popularity in novel-view scene synthesis. It addresses the challenges of lengthy training times and slow rendering speeds associated with Neural Radiance Fields (NeRFs). Through rapid, differentiable rasterization of 3D Gaussians, 3D-GS achieves re…

2024

Enhancing Object Grasping Efficiency with Deep Learning and Post-processing for Multi-finger Robotic Hands

IROS 2024poster

This paper builds upon the well-established ML-based grasping technique, known as the Grasp-Rectangle (GR) method. The original GR method made two simplifying assumptions: it was designed exclusively for two-finger grippers, and it assumed that the gripper would approach objects solely from a top-do…

Cited by 0SourceScholar
2024

Investigating Style Similarity in Diffusion Models

ECCV 2024poster

"Generative models are now widely used by graphic designers and artists. Prior works have shown that these models remember and often replicate content from their training data during generation. Hence as their proliferation increases, it has become important to perform a database search to determine…

Cited by 0SourcePDFScholar
2024

LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation

ECCV 2024poster

"Neural Radiance Fields (NeRFs) have revolutionized the reconstruction of static scenes and objects in 3D, offering unprecedented quality. However, extending NeRFs to model dynamic objects or object articulations remains a challenging problem. Previous works have tackled this issue by focusing on pa…

Cited by 7SourcePDFScholar
2024

LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors

ECCV 2024poster

"We present a simple self-supervised method to enhance the performance of ViT features for dense downstream tasks. Our Lightweight Feature Transform (LiFT) is a straightforward and compact postprocessing network that can be applied to enhance the features of any pre-trained ViT backbone. LiFT is fas…

Cited by 5SourcePDFScholar
2024

Robust Partitioned Visual Servoing for Aerial Manipulation Utilizing Controllable-space Image Planning and Adaptive Image Representation

IROS 2024poster

In the pursuit of object retrieval using an aerial manipulator, developing robust visual servoing techniques in the presence of projection and motion model uncertainties is paramount. This paper proposes a novel approach to conducting image-space planning within the controllable-space of the aerial…

Cited by 0SourceScholar
2023

ASIC: Aligning Sparse in-the-wild Image Collections

ICCV 2023oral

We present a method for joint alignment of sparse in-the-wild image collections of an object category. Most prior works assume either ground-truth keypoint annotations or a large dataset of images of a single object category. However, neither of the above assumptions hold true for the long-tail of t…

Cited by 21PDFcodeScholar
2023

Chop & Learn: Recognizing and Generating Object-State Compositions

ICCV 2023poster

Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resulting object state changes. We propose a new benchmark suite Chop & Learn, to acc…

Cited by 18PDFcodeScholar
2023

LilNetX: Lightweight Networks with EXtreme Model Compression and Structured Sparsification

ICLR 2023poster

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach these problems one at a time and often require post-processing or multistage training which become less practical and do n…

2023

SHACIRA: Scalable HAsh-grid Compression for Implicit Neural Representations

ICCV 2023poster

Implicit Neural Representations (INR) or neural fields have emerged as a popular framework to encode multimedia signals such as images and radiance fields while retaining high-quality. Recently, learnable feature grids such as Instant-NGP have allowed significant speed-up in the training as well as…

Cited by 30PDFcodeScholar
2023

Teaching Matters: Investigating the Role of Supervision in Vision Transformers

CVPR 2023poster

Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradigms is not well explored. We compare ViTs trained through different methods of supervision, and show that they learn a di…

2022

Neural-Guided Runtime Prediction of Planners for Improved Motion and Task Planning with Graph Neural Networks

IROS 2022poster

The past decade has amply demonstrated the remarkable functionality that can be realized by learning complex input/output relationships. Algorithmically, one of the most important and opaque relationships is that between a problem's structure and an effective solution method. Here, we quantitatively…

Cited by 5SourceScholar
2021

LayoutTransformer: Layout Generation and Completion With Self-Attention

ICCV 2021poster

We address the problem of scene layout generation for diverse domains such as images, mobile applications, documents, and 3D objects. Most complex scenes, natural or human-designed, can be expressed as a meaningful arrangement of simpler compositional graphical primitives. Generating a new layout or…

Cited by 184PDFcodeScholar
2021

PatchGame: Learning to Signal Mid-level Patches in Referential Games

NeurIPS 2021poster

We study a referential game (a type of signaling game) where two agents communicate with each other via a discrete bottleneck to achieve a common goal. In our referential game, the goal of the speaker is to compose a message or a symbolic representation of "important" image patches, while the task f…

2021

The Lottery Ticket Hypothesis for Object Recognition

CVPR 2021poster

Recognition tasks, such as object recognition and keypoint estimation, have seen widespread adoption in recent years. Most state-of-the-art methods for these tasks use deep networks that are computationally expensive and have huge memory footprints. This makes it exceedingly difficult to deploy thes…

Cited by 76PDFcodeScholar
2021

Toward Observation Based Least Restrictive Collision Avoidance Using Deep Meta Reinforcement Learning

RA-L 2021

This letter presents the Observation-based Least-Restrictive Collision Avoidance Module (OLR-CAM) that can be added to any autonomous robot working in a shared environment and provide a high-level safety layer to the existing policy for each robot. The OLR-CAM takes raw sensory observations as input

Cited by 4SourceScholar
2019

Towards an Integrated Autonomous Data-Driven Grasping System with a Mobile Manipulator

ICRA 2019poster

We present an integrated grasping system for a mobile manipulator to grasp an unknown object of interest (OI) in an unknown environment. The system autonomously scans its environment, models the OI, plans and executes a grasp, while taking into account base pose uncertainty. Due to inherent line of…

Cited by 16SourceScholar