← Search

Krishna Somandepalli

10 accepted papers

2024

A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation

NeurIPS 2024poster

Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we…

2024

VideoPoet: A Large Language Model for Zero-Shot Video Generation

ICML 2024oral

We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs -- including images, videos, text, and audio. The training protocol follows that…

Cited by 257SourcePDFScholar
2023

A Dataset for Audio-Visual Sound Event Detection in Movies

ICASSP 2023accepted

Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at s…

Cited by 0SourceScholar
2023

Contextually-Rich Human Affect Perception Using Multimodal Scene Information

ICASSP 2023accepted

The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived…

Cited by 0SourceScholar
2023

Heterogeneous Graph Learning for Acoustic Event Classification

ICASSP 2023accepted

Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, graph structure does not appear naturally in audiovisual data. Graphs for audiovis…

Cited by 0SourceScholar
2020

Robust Speaker Recognition Using Unsupervised Adversarial Invariance

ICASSP 2020accepted

In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker emb…

Cited by 0SourceScholar
2020

Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context

ICASSP 2020accepted

Due to its ability to visualize and measure the dynamics of vocal tract shaping during speech production, real-time magnetic resonance imaging (rtMRI) has emerged as one of the prominent research tools. The ability to track different articulators such as the tongue, lips, velum, and the pharynx is a…

Cited by 8SourceScholar
2019

Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies

ICASSP 2019accepted

The ability to robustly cluster faces in movies is a necessary step in understanding media content representations of people along dimensions such as gender and age. Building upon the successes of sparse subspace clustering (SSC) in uncovering the underlying structure of the data, in this paper we p…

Cited by 0SourceScholar
2019

Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation

ICASSP 2019accepted

Speech activity detection in highly variable acoustic conditions is a challenging task. Many approaches to detect speech activity in such conditions involve an inherent knowledge of the noise types involved. Movie audio can offer an excellent research test-bed for developing speech activity models.…

Cited by 0SourceScholar
2019

Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders

ICASSP 2019accepted

Audio-signal acquisition as part of wearable sensing adds an important dimension for applications such as understanding human behaviors. As part of a large study on work place behaviours, we collected audio data from individual hospital staff using custom wearable recorders. The audio features colle…

Cited by 0SourceScholar