← Search

Ross Cutler

20 accepted papers

2026

Human-in-the-Loop Bandwidth Estimation for Quality of Experience Optimization in Real-Time Video Communication

AAAI 2026technical

The quality of experience (QoE) delivered by video conferencing systems is significantly influenced by accurately estimating the time-varying available bandwidth between the sender and receiver. Bandwidth estimation for real-time communications remains an open challenge due to rapidly evolving netwo

Cited by 0SourcePDFScholar
2024

A Real-Time Active Speaker Detection System Integrating an Audio-Visual Signal with a Spatial Querying Mechanism

ICASSP 2024accepted

We introduce a distinctive real-time, causal, neural network-based active speaker detection system optimized for low-power edge computing. This system drives a virtual cinematography module and is deployed on a commercial device. The system uses data originating from a microphone array and a 360-deg…

Cited by 0SourceScholar
2024

Topic-Conversation Relevance (TCR) Dataset and Benchmarks

NeurIPS 2024poster

Workplace meetings are vital to organizational collaboration, yet a large percentage of meetings are rated as ineffective. To help improve meeting effectiveness by understanding if the conversation is on topic, we create a comprehensive Topic-Conversation Relevance (TCR) dataset that covers a variet…

2024

VCD: A Video Conferencing Dataset for Video Compression

ICASSP 2024accepted

Commonly used datasets for evaluating video codecs are all very high quality and not representative of video typically used in video conferencing scenarios. We present the Video Conferencing Dataset (VCD) for evaluating video codecs for real-time communication, the first such dataset focused on vide…

Cited by 0SourceScholar
2023

AURA: Privacy-Preserving Augmentation to Improve Test Set Diversity in Speech Enhancement

ICASSP 2023accepted

Speech enhancement models running in production environments are commonly trained on publicly available data. This approach leads to regressions due to the lack of training/testing on representative customer data. Moreover, due to privacy reasons, developers cannot listen to customer content. This ‘…

Cited by 0SourceScholar
2023

LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing Calls

ICASSP 2023accepted

Current state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that occur during videoconferencing calls. In this paper, we pres…

Cited by 0SourceScholar
2023

Real-Time Speech Interruption Analysis: from Cloud to Client Deployment

ICASSP 2023accepted

Meetings are an essential form of communication for all types of organizations, and remote collaboration systems have been much more widely used since the COVID-19 pandemic. One major issue with remote meetings is that it is challenging for remote participants to interrupt and speak. We have recentl…

Cited by 0SourceScholar
2022

AECMOS: A Speech Quality Assessment Metric for Echo Impairment

ICASSP 2022accepted

Traditionally, the quality of acoustic echo cancellers is evaluated using intrusive speech quality assessment measures such as ERLE [1] and PESQ [2], or by carrying out subjective laboratory tests [3], [4]. Unfortunately, the former are not well correlated with human subjective measures, while the l…

Cited by 60SourceScholar
2022

Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors

ICASSP 2022accepted

Human subjective evaluation is the "gold standard" to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. We have recently developed a non-intrusive speech quality metric called Deep Noise Suppression Mean Opinion Score (DNSMOS…

Cited by 0SourceScholar
2022

ICASSP 2022 Acoustic Echo Cancellation Challenge

ICASSP 2022accepted

The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios,…

Cited by 83SourceScholar
2022

Icassp 2022 Deep Noise Suppression Challenge

ICASSP 2022accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-sourc…

Cited by 0SourceScholar
2021

Crowdsourcing Approach for Subjective Evaluation of Echo Impairment

ICASSP 2021accepted

The quality of acoustic echo cancellers (AECs) in real-time communication systems is typically evaluated using objective metrics like ERLE [1] and PESQ [2], and less commonly with lab-based subjective tests like ITU-T Rec. P.831 [3]. We will show that these objective measures are not well correlated…

Cited by 0SourceScholar
2021

Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors

ICASSP 2021accepted

Human subjective evaluation is the "gold standard" to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. The conventional and widely used metrics require a reference clean speech signal, which is unavailable in real recordings…

Cited by 0SourceScholar
2021

ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results

ICASSP 2021accepted

The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performa…

Cited by 0SourceScholar
2021

ICASSP 2021 Deep Noise Suppression Challenge

ICASSP 2021accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to tr…

Cited by 0SourceScholar
2020

Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing

ICASSP 2020accepted

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a…

Cited by 0SourceScholar
2020

Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement

ICASSP 2020accepted

This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, sing…

Cited by 0SourceScholar
2019

Non-intrusive Speech Quality Assessment Using Neural Networks

ICASSP 2019accepted

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neura…

Cited by 0SourceScholar