← Search

Sauradip Nag

10 accepted papers

2025

OmniCount: Multi-label Object Counting with Semantic-Geometric Priors

AAAI 2025technical

Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, these strategies come with their own set of limitations, such as the need for ma…

Cited by 2SourcePDFScholar
2025

RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation

NeurIPS 2025poster

The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this w…

Cited by 0SourcecodeScholar
2025

SMITE: Segment Me In TimE

ICLR 2025poster

Segmenting an object in a video presents significant challenges. Each pixel must be accurately labeled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks…

2024

DiffSED: Sound Event Detection with Denoising Diffusion

AAAI 2024technical

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the split-and-classify (i.e., frame-level) strategy or the more principled event-level modeling approach, all existing methods…

2023

DiffTAD: Temporal Action Detection with Proposal Denoising Diffusion

ICCV 2023poster

We propose a new formulation of temporal action detection (TAD) with denoising diffusion, DiffTAD in short. Taking as input random temporal proposals, it can yield action proposals accurately given an untrimmed long video. This presents a generative modeling perspective, against previous discriminat…

Cited by 43PDFcodeScholar
2022

Proposal-Free Temporal Action Detection via Global Segmentation Mask Learning

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on generating an overwhelmingly large number of proposals per video. This leads to complex model designs due to proposal generation and/or per-proposal action instance evaluation and the resultant high computational cost. In this work, for the f…

2022

Semi-Supervised Temporal Action Detection with Proposal-Free Masking

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on a large number of training data with segment-level annotations. Collecting and annotating such a training set is thus highly expensive and unscalable. Semi-supervised TAD (SS-TAD) alleviates this problem by leveraging unlabeled videos freely…

2022

Zero-Shot Temporal Action Detection via Vision-Language Prompting

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on large training data including segment-level annotations, limited to recognizing previously seen classes alone during inference. Collecting and annotating a large training set for each class of interest is costly and hence unscalable. Zero-sho…

2019

Facial Micro-expression Spotting and Recognition Using Time Contrasted Feature with Visual Memory

ICASSP 2019accepted

Facial micro-expressions are sudden involuntary minute muscle movements which reveal true emotions that people try to conceal. Spotting a micro-expression and recognizing it is a major challenge owing to its short duration and intensity. Many works pursued traditional and deep learning based approac…

Cited by 0SourceScholar