← Search

Onkar Kishor Susladkar

3 accepted papers

2025

MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion

ICLR 2025spotlight

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we introduce the 3D Mobile Inverted Vector-Quantization Variat…

2024

Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS.

EMNLP 2024finding

This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech in the Bahasa language. The dataset, spanning 55.00 hours and 52K audio recordings, integrates diverse textual sources, e…

Cited by 0SourcePDFScholar
2024

GRIZAL: Generative Prior-guided Zero-Shot Temporal Action Localization

EMNLP 2024main

Zero-shot temporal action localization (TAL) aims to temporally localize actions in videos without prior training examples. To address the challenges of TAL, we offer GRIZAL, a model that uses multimodal embeddings and dynamic motion cues to localize actions effectively. GRIZAL achieves sample diver…

Cited by 0SourcePDFScholar