← Search

Aniruddha Mahapatra

6 accepted papers

2025

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces

ICCV 2025poster

Video tokenizers are essential for latent video diffusion models, converting raw video data into spatiotemporally compressed latent spaces for efficient training. However, extending state-of-the-art video tokenizers to achieve a temporal compression ratio beyond 4x without increasing channel capacit…

2025

REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder

ICCV 2025poster

We present a novel perspective on learning video embedders for generative modeling: rather than requiring an exact reproduction of an input video, an effective embedder should focus on synthesizing visually plausible reconstructions. This relaxed criterion enables substantial improvements in compres…

Cited by 0SourcePDFScholar
2024

On the Content Bias in Frechet Video Distance

CVPR 2024poster

Frechet Video Distance (FVD) a prominent metric for evaluating video generation models is known to conflict with human perception occasionally. In this paper we aim to explore the extent of FVD's bias toward frame quality over temporal realism and identify its sources. We first quantify the FVD's se…

Cited by 35SourcePDFScholar
2022

Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language Models

EMNLP 2022main

Transformer-based language models trained on large natural language corpora have been very useful in downstream entity extraction tasks. However, they often result in poor performances when applied to domains that are different from those they are pretrained on. Continued pretraining using unlabeled…

2021

SemIE: Semantically-Aware Image Extrapolation

ICCV 2021poster

We propose a semantically-aware novel paradigm to perform image extrapolation that enables the addition of new object instances. All previous methods are limited in their capability of extrapolation to merely extending the already existing objects in the image. However, our proposed approach focuses…

Cited by 16PDFScholar