← Search

Amit Kumar Singh Yadav

5 accepted papers

2025

DiffSSD: A Diffusion-Based Dataset For Speech Forensics

ICASSP 2025accepted

Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do no…

Cited by 0SourceScholar
2025

ItDPDM: Information-Theoretic Discrete Poisson Diffusion Model

NeurIPS 2025poster

Generative modeling of non-negative, discrete data, such as symbolic music, remains challenging due to two persistent limitations in existing methods. Firstly, many approaches rely on modeling continuous embeddings, which is suboptimal for inherently discrete data distributions. Secondly, most model…

Cited by 0SourceScholar
2025

Speech-N-LlaMA: Improving Speech LLMs with Multi-Pass Training

ICASSP 2025accepted

Speech LLMs use speech embeddings as the prompt to a Large Language Model (LLM) and generate human readable text for the speech signal in an autoregressive manner. Teacher-forcing is a common approach used for training Speech LLMs, which is dissimilar to the procedure used during inference, creating…

Cited by 0SourceScholar
2024

Mdrt: Multi-Domain Synthetic Speech Localization

ICASSP 2024accepted

With recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many meth…

Cited by 0SourceScholar
2023

ASSD: Synthetic Speech Detection in the AAC Compressed Domain

ICASSP 2023accepted

Synthetic human speech signals have become very easy to generate given modern text-to-speech methods. When these signals are shared on social media they are often compressed using the Advanced Audio Coding (AAC) standard. Our goal is to study if a small set of coding metadata contained in the AAC co…

Cited by 0SourceScholar