← Search

Chandrashekhar Lavania

7 accepted papers

2025

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

ICASSP 2025accepted

Audio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in du…

Cited by 0SourceScholar
2024

Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers’ Opinion Scores

ECCV 2024poster

"Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally accepted evaluation metrics also play an important role in advanc…

Cited by 1SourcePDFScholar
2023

Multi-Scale Compositional Constraints for Representation Learning on Videos

ICASSP 2023accepted

Combining simple concepts to form structured thoughts and decomposing complex concepts into their constituents is one key characteristic of human cognition. In this work we extract video representations by combining multi-scale processing with compositional constraints, i.e., we constrain the latent…

Cited by 0SourceScholar
2022

Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation

ICASSP 2022accepted

Audio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video…

Cited by 0SourceScholar
2019

Fixing Mini-batch Sequences with Hierarchical Robust Partitioning

AISTATS 2019poster

We propose a general and efficient hierarchical robust partitioning framework to generate a deterministic sequence of mini-batches, one that offers assurances of being high quality, unlike a randomly drawn sequence. We compare our deterministically generated mini-batch sequences to randomly generat…

Cited by 12SourcePDFScholar
2017

Reducing total latency in online real-time inference and decoding via combined context window and model smoothing latencies

ICASSP 2017accepted

Real-time low-latency online inference and decoding in sequential probabilistic models are important in many interactive systems, including automatic speech recognition (ASR) and streaming environments. We study total inference latency (TL) in such systems, the additively combined latency of the inh…

Cited by 0SourceScholar