← Search

Sharath Girish

9 accepted papers

2026

AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation

CVPR 2026

Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing methods lack fine-grained temporal control over subject appearance and disappearance, which are essential for applicatio

Cited by 0SourceScholar
2025

Improving the Diffusability of Autoencoders

ICML 2025poster

Latent diffusion models have emerged as the leading approach for generating high-quality images and videos, utilizing compressed latent representations to reduce the computational burden of the diffusion process. While recent advancements have primarily focused on scaling diffusion backbones and imp…

2024

QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos

NeurIPS 2024poster

Online free-viewpoint video (FVV) streaming is a challenging problem, which is relatively under-explored. It requires incremental on-the-fly updates to a volumetric representation, fast training and rendering to satisfy realtime constraints and a small memory footprint for efficient transmission. If…

Cited by 0SourcePDFScholar
2023

LilNetX: Lightweight Networks with EXtreme Model Compression and Structured Sparsification

ICLR 2023poster

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach these problems one at a time and often require post-processing or multistage training which become less practical and do n…

2023

NIRVANA: Neural Implicit Representations of Videos With Adaptive Networks and Autoregressive Patch-Wise Modeling

CVPR 2023poster

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architec…

2023

SHACIRA: Scalable HAsh-grid Compression for Implicit Neural Representations

ICCV 2023poster

Implicit Neural Representations (INR) or neural fields have emerged as a popular framework to encode multimedia signals such as images and radiance fields while retaining high-quality. Recently, learnable feature grids such as Instant-NGP have allowed significant speed-up in the training as well as…

Cited by 30PDFcodeScholar
2021

The Lottery Ticket Hypothesis for Object Recognition

CVPR 2021poster

Recognition tasks, such as object recognition and keypoint estimation, have seen widespread adoption in recent years. Most state-of-the-art methods for these tasks use deep networks that are computationally expensive and have huge memory footprints. This makes it exceedingly difficult to deploy thes…

Cited by 76PDFcodeScholar
2021

Towards Discovery and Attribution of Open-World GAN Generated Images

ICCV 2021poster

With the recent progress in Generative Adversarial Networks (GANs), it is imperative for media and visual forensics to develop detectors which can identify and attribute images to the model generating them. Existing works have shown to attribute images to their corresponding GAN sources with high ac…

Cited by 71PDFcodeScholar