← Search

Bharat Singh

14 accepted papers

2026

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

CVPR 2026

Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the strong coupling between conditional and unconditional branches in standard classifier-free guidance. We introduce a training-free framework that en

Cited by 0SourceScholar
2023

LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation

ICCV 2023oral

Large-scale pre-training tasks like image classification, captioning, or self-supervised techniques do not incentivize learning the semantic boundaries of objects. However, recent generative foundation models built using text-based latent diffusion techniques may learn semantic boundaries. This is b…

Cited by 38PDFcodeScholar
2018

R-FCN-3000 at 30fps: Decoupling Detection and Classification

CVPR 2018poster

We propose a modular approach towards large-scale real-time object detection by decoupling objectness detection and classification. We exploit the fact that many object classes are visually similar and share parts. Thus, a universal objectness detector can be learned for class-agnostic object detect…

2017

Soft-NMS -- Improving Object Detection With One Line of Code

ICCV 2017poster

Non-maximum suppression is an integral part of the object detection pipeline. First, it sorts all detection boxes on the basis of their scores. The detection box M with the maximum score is selected and all other detection boxes with a significant overlap (using a pre-defined threshold) with M are s…

Cited by 2109PDFScholar
2017

Son of Zorn's lemma: Targeted style transfer using instance-aware semantic segmentation

ICASSP 2017accepted

Style transfer is an important task in which the style of a source image is mapped onto that of a target image. The method is useful for synthesizing derivative works of a particular artist or specific painting. This work considers targeted style transfer, in which the style of a template image is u…

Cited by 0SourceScholar
2017

Temporal Context Network for Activity Localization in Videos

ICCV 2017poster

We present a Temporal Context Network (TCN) for precise temporal localization of human activities. Similar to the Faster-RCNN architecture, proposals are placed at equal intervals in a video which span multiple temporal scales. We propose a novel representation for ranking these proposals. Since poo…

Cited by 318PDFScholar
2016

A Multi-Stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection

CVPR 2016poster

We present a multi-stream bi-directional recurrent neural network for fine-grained action detection. Recently, two-stream convolutional neural networks (CNNs) trained on stacked optical flow and image frames have been successful for action recognition in videos. Our system uses a tracking algorithm…

Cited by 606PDFScholar
2016

Training Neural Networks Without Gradients: A Scalable ADMM Approach

ICML 2016poster

With the growing importance of large network models and enormous training datasets, GPUs have become increasingly necessary to train neural networks. This is largely because conventional optimization algorithms rely on stochastic gradient methods that don’t scale well to large numbers of cores in a…

Cited by 341SourcePDFScholar
2015

Selecting Relevant Web Trained Concepts for Automated Event Retrieval

ICCV 2015poster

Complex event retrieval is a challenging research problem, especially when no training videos are available. An alternative to collecting training videos is to train a large semantic concept bank a priori. Given a text description of an event, event retrieval is performed by selecting concepts lingu…

Cited by 42PDFScholar