← Search

Sujoy Paul

14 accepted papers

2025

Masked Generative Nested Transformers with Decode Time Scaling

ICML 2025poster

Recent advances in visual generation have made significant strides in producing content of exceptional quality. However, most methods suffer from a fundamental problem - a bottleneck of inference computational efficiency. Most of these algorithms involve multiple passes over a transformer model to g…

Cited by 0SourcePDFScholar
2024

Mixture of Nested Experts: Adaptive Processing of Visual Tokens

NeurIPS 2024poster

The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based models scale effectively to large data regimes, they fail to capitalize on this inher…

Cited by 8SourcePDFScholar
2023

Test-time Adaptation with Slot-Centric Models

ICML 2023poster

Current visual detectors, though impressive within their training distribution, often fail to parse out-of-distribution scenes into their constituent entities. Recent test-time adaptation methods use auxiliary self-supervised losses to adapt the network parameters to each test example independently…

2022

Novel Class Discovery without Forgetting

ECCV 2022poster

"Humans possess an innate ability to identify and differentiate instances that they are not familiar with, by leveraging and adapting the knowledge that they have acquired so far. Importantly, they achieve this without deteriorating the performance on their earlier learning. Inspired by this, we ide…

Cited by 52SourcePDFScholar
2021

Cross-domain Imitation from Observations

ICML 2021oral

Imitation learning seeks to circumvent the difficulty in designing proper reward functions for training agents by utilizing expert behavior. With environments modeled as Markov Decision Processes (MDP), most of the existing imitation algorithms are contingent on the availability of expert demonstrat…

Cited by 45SourcePDFScholar
2021

Unsupervised Multi-Source Domain Adaptation Without Access to Source Data

CVPR 2021poster

Unsupervised Domain Adaptation (UDA) aims to learn a predictor model for an unlabeled dataset by transferring knowledge from a labeled source data, which has been trained on similar tasks. However, most of these conventional UDA approaches have a strong assumption of having access to the source data…

Cited by 200PDFScholar
2020

Domain Adaptive Semantic Segmentation Using Weak Labels

ECCV 2020poster

We propose a novel framework for domain adaptation in semantic segmentation with image-level weak labels in the target domain. The weak labels may be obtained based on a model prediction for unsupervised domain adaptation (UDA), or from a human oracle in a new weakly-supervised domain adaptation (WD…

Cited by 96SourcePDFScholar
2019

Weakly Supervised Video Moment Retrieval From Text Queries

CVPR 2019poster

There have been a few recent methods proposed in text to video moment retrieval using natural language queries, but requiring full supervision during training. However, acquiring a large number of training videos with temporal boundary annotations for each text description is extremely time-consumin…

Cited by 238PDFcodeScholar
2018

Exploiting Transitivity for Learning Person Re-Identification Models on a Budget

CVPR 2018poster

Minimization of labeling effort for person re-identification in camera networks is an important problem as most of the existing popular methods are supervised and they require large amount of manual annotations, acquiring which is a tedious job. In this work, we focus on this labeling effort minimiz…

Cited by 22SourcePDFScholar
2018

Incorporating Scalability in Unsupervised Spatio- Temporal Feature Learning

ICASSP 2018accepted

Deep neural networks are efficient learning machines which leverage upon a large amount of manually labeled data for learning discriminative features. However, acquiring substantial amount of supervised data, especially for videos can be a tedious job across various computer vision tasks. This neces…

Cited by 0SourceScholar
2018

W-TALC: Weakly-supervised Temporal Activity Localization and Classification

ECCV 2018poster

Most activity localization methods in the literature suffer from the burden of frame-wise annotation requirement. Learning from weak labels may be a potential solution towards reducing such manual labeling effort. Recent years have witnessed a substantial influx of tagged videos on the Internet, whi…

2017

The Impact of Typicality for Informative Representative Selection

CVPR 2017poster

In computer vision, selection of the most informative samples from a huge pool of training data in order to learn a good recognition model is an active research problem. Furthermore, it is also useful to reduce the annotation cost, as it is time consuming to annotate unlabeled samples. In this paper…

Cited by 10PDFScholar