← Search

Kaustav Kundu

13 accepted papers

2024

Robustness Preserving Fine-tuning using Neuron Importance

ECCV 2024poster

"Robust fine-tuning aims to adapt a vision-language model to downstream tasks while preserving its zero-shot capabilities on unseen data. Recent studies have introduced fine-tuning strategies to improve in-distribution (ID) performance on the downstream tasks while minimizing deterioration in out-of…

Cited by 0SourcePDFScholar
2022

Hierarchical Self-Supervised Representation Learning for Movie Understanding

CVPR 2022poster

Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understanding and propose a novel hierarchical self-supervised pretraining strategy that separately pretrains each level of our h…

Cited by 29PDFcodeScholar
2022

TubeR: Tubelet Transformer for Video Action Detection

CVPR 2022oral

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in video by simulta…

Cited by 95PDFScholar
2022

What To Look at and Where: Semantic and Spatial Refined Transformer for Detecting Human-Object Interactions

CVPR 2022oral

We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize humans and objects, and predicts their interactions. Differently from previous Transformer-based HOI approaches, which mostl…

Cited by 66PDFcodeScholar
2021

Positive-Congruent Training: Towards Regression-Free Model Updates

CVPR 2021poster

Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correc…

Cited by 63PDFScholar
2018

SurfConv: Bridging 3D and 2D Convolution for RGBD Images

CVPR 2018poster

The last few years have seen approaches trying to combine the increasing popularity of depth sensors and the success of the convolutional neural networks. Using depth as additional channel alongside the RGB input has the scale variance problem present in image convolution based approaches. On the ot…

2016

Monocular 3D Object Detection for Autonomous Driving

CVPR 2016poster

The goal of this paper is to perform 3D object detection in single monocular images in the domain of autonomous driving. Our method first aims to generate a set of candidate class-specific object proposals, which are then run through a standard CNN pipeline to obtain high-quality object detections.…

Cited by 1263PDFScholar
2015

3D Object Proposals for Accurate Object Class Detection

NeurIPS 2015poster

The goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane a…

Cited by 1092SourcePDFScholar
2015

Rent3D: Floor-Plan Priors for Monocular Layout Estimation

CVPR 2015poster

The goal of this paper is to enable a 3D "virtual-tour" of an apartment given a small set of monocular images of different rooms, as well as a 2D floor plan. We frame the problem as inference in a Markov Random Field which reasons about the layout of each room and its relative pose (3D rotation and…

Cited by 177SourcePDFScholar