← Search

Vimal Bhat

10 accepted papers

2025

Beyond Speaker Identity: Text Guided Target Speech Extraction

ICASSP 2025accepted

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in…

Cited by 0SourceScholar
2025

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

ACL 2025finding

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…

Cited by 0SourcePDFScholar
2025

Detect, Disambiguate, and Translate: On-Demand Visual Reasoning for Multimodal Machine Translation with Large Vision-Language Models

NAACL 2025long

Multimodal machine translation (MMT) aims to leverage additional modalities to assist in language translation. With limited parallel data, current MMT systems rely heavily on monolingual English captioning data. These systems face three key issues: they often overlook that visual signals are unneces…

Cited by 0SourcePDFScholar
2025

Learning Rich Speech Representations with Acoustic-Semantic Factorization

ICASSP 2025accepted

Self-supervised pretraining has transformed speech representation learning, enabling models to generalize across various downstream tasks. However, empirical studies have highlighted two notable gaps. First, different speech tasks require varying levels of acoustic and semantic information, which ar…

Cited by 0SourceScholar
2025

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

NeurIPS 2025poster

Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens or hundreds of tokens, the limited context length of large language models (LLMs) forces the VLMs to perceive the frames…

Cited by 0SourceScholar
2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

EMNLP 2025

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X util

Cited by 0SourcePDFScholar
2023

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

ICCV 2023poster

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (>60min). In this paper, we introduce Mu…

Cited by 4PDFcodeScholar
2023

Motion-Guided Masking for Spatiotemporal Representation Learning

ICCV 2023poster

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video understanding. This suggests that the random masking strategy that…

Cited by 16PDFScholar
2021

Shot Contrastive Self-Supervised Learning for Scene Boundary Detection

CVPR 2021poster

Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present…

Cited by 87PDFScholar