← Search

Tanaya Guha

14 accepted papers

2025

Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions

ICASSP 2025accepted

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener’s response to the interlocutor’s speech. Although significant progress has been made in the context of generating co-speech gestures, generating listener’s r…

Cited by 0SourceScholar
2023

Heterogeneous Graph Learning for Acoustic Event Classification

ICASSP 2023accepted

Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, graph structure does not appear naturally in audiovisual data. Graphs for audiovis…

Cited by 0SourceScholar
2022

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

ECCV 2022poster

"Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a novel spatial-temporal graph learning framework that can solv…

2021

In Defense of Scene Graphs for Image Captioning

ICCV 2021poster

The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics such as object entities, relationships and…

Cited by 53PDFcodeScholar
2019

Computational Analysis of Gaze Behavior in Autism During Interaction with Virtual Agents

ICASSP 2019accepted

Individuals with Autism spectrum disorder (ASD) are known to have significantly impaired social interaction and communication abilities. These impairments are characterized by their difficulties in using and perceiving non-verbal cues, such as facial expressions. The difficulty in processing communi…

Cited by 0SourceScholar
2016

A multimodal mixture-of-experts model for dynamic emotion prediction in movies

ICASSP 2016accepted

This paper addresses the problem of continuous emotion prediction in movies from multimodal cues. The rich emotion content in movies is inherently multimodal, where emotion is evoked through both audio (music, speech) and video modalities. To capture such affective information, we put forth a set of…

Cited by 0SourceScholar
2016

Opening big in box office? Trailer content can help

ICASSP 2016accepted

Computational prediction of a movie's financial success usually relies only on metadata such as - genre, budget, actors, Motion Picture Association of America (MPAA) rating and critics' reviews. We argue that movie trailers, created to invoke viewers' interest and curiosity about a movie, carry comp…

Cited by 0SourceScholar
2015

Computationally deconstructing movie narratives: An informatics approach

ICASSP 2015accepted

In general, popular films and screenplays follow a well defined storytelling paradigm that comprises three essential segments or acts: exposition (act I), conflict (act II) and resolution (act III). Deconstructing a movie into its narrative units can enrich semantic understanding of movies, and help…

Cited by 0SourceScholar
2015

On quantifying facial expression-related atypicality of children with Autism Spectrum Disorder

ICASSP 2015accepted

by adult observers. This paper focuses on data driven ways to analyze and quantify atypicality in facial expressions of children with ASD. Our objective is to uncover those characteristics of facial gestures that induce the sense of perceived atypicality in observers. Using a carefully collected mot…

Cited by 0SourceScholar