← Search

Julien Epps

12 accepted papers

2025

Rethinking Mamba in Speech Processing by Self-Supervised Models

ICASSP 2025accepted

The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enh…

Cited by 0SourceScholar
2025

SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information

ACL 2025finding

Large Language Models (LLMs) have been increasingly adopted for health-related tasks, yet their performance in depression detection remains limited when relying solely on text input. While Retrieval-Augmented Generation (RAG) typically enhances LLM capabilities, our experiments indicate that traditi…

Cited by 0SourcePDFScholar
2024

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

EMNLP 2024main

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in healthcare applications. However, the application of LLMs in the identification and a…

Cited by 11SourcePDFScholar
2021

Automatic Elicitation Compliance for Short-Duration Speech Based Depression Detection

ICASSP 2021accepted

Detecting depression from the voice in naturalistic environments is challenging, particularly for short-duration audio recordings. This enhances the need to interpret and make optimal use of elicited speech. The rapid consonant-vowel syllable combination ‘pataka’ has frequently been selected as a cl…

Cited by 0SourceScholar
2020

Exploiting Vocal Tract Coordination Using Dilated CNNS For Depression Detection In Naturalistic Environments

ICASSP 2020accepted

Depression detection from speech continues to attract significant research attention but remains a major challenge, particularly when the speech is acquired from diverse smartphones in natural environments. Analysis methods based on vocal tract coordination have shown great promise in depression and…

Cited by 0SourceScholar
2019

Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection

ICASSP 2019accepted

The security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. T…

Cited by 0SourceScholar
2019

Evaluation Measures for Depression Prediction and Affective Computing

ICASSP 2019accepted

A variety of evaluation measures are being used to validate systems in depression prediction and affective computing. Among them, the most common measures focus on the error between the ground truth and predictions. However, when the ground truth is ordinal such as in psychiatric scores, ranking inf…

Cited by 0SourceScholar
2019

Speech Landmark Bigrams for Depression Detection from Naturalistic Smartphone Speech

ICASSP 2019accepted

Detection of depression from speech has attracted significant research attention in recent years but remains a challenge, particularly for speech from diverse smartphones in natural environments. This paper proposes two sets of novel features based on speech landmark bigrams associated with abrupt s…

Cited by 0SourceScholar
2019

Transmission Line Cochlear Model Based AM-FM Features for Replay Attack Detection

ICASSP 2019accepted

This paper focuses on providing a countermeasure to replay attack which is the simplest and more accessible form of attack used to spoof automatic speaker verification systems. Specifically, it proposes the use of the transmission line cochlear model, which resembles the human cochlea more accuratel…

Cited by 13SourceScholar
2017

A PLLR and multi-stage Staircase Regression framework for speech-based emotion prediction

ICASSP 2017accepted

Continuous prediction of dimensional emotions (e.g. arousal and valence) has attracted increasing research interest recently. When processing emotional speech signals, phonetic features have been rarely used due to the assumption that phonetic variability is a confounding factor that degrades emotio…

Cited by 0SourceScholar
2015

Weighted pairwise Gaussian likelihood regression for depression score prediction

ICASSP 2015accepted

This paper presents a technique in which feature vectors are mapped onto ordinal ranges of clinical depression scores using weighted pairwise Gaussians. The position of a test vector with respect to these partitions is used to perform depression score prediction. Results found on a set of spectral a…

Cited by 0SourceScholar