← Search

Javier Hernando

6 accepted papers

2026

REVISITING DIRECT SPEECH-TO-TEXT TRANSLATION WITH SPEECH LLMS: BETTER SCALING THAN COT PROMPTING?

ICASSP 2026poster

Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then translate it. CoT typically outperforms direct prompting primarily because it can…

Cited by 0SourcePDFScholar
2024

Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge

ACL 2024findings

Recent research has explored methods for updating and modifying factual knowledge in large language models, often focusing on specific multi-layer perceptron blocks. This study expands on this work by examining the effectiveness of existing knowledge editing methods across languages and delving into…

2020

I-Vector Transformation Using K-Nearest Neighbors for Speaker Verification

ICASSP 2020accepted

Probabilistic Linear Discriminant Analysis (PLDA) is the most efficient backend for i-vectors. However, it requires labeled background data which can be difficult to access in practice. Unlike PLDA, cosine scoring avoids speaker-labels at the cost of degrading the performance. In this work, we propo…

Cited by 0SourceScholar
2016

Work-efficient parallel non-maximum suppression for embedded GPU architectures

ICASSP 2016accepted

With the emergence of GPU computing, deep neural networks have become a widely used technique for advancing research in the field of image and speech processing. In the context of object and event detection, sliding-window classifiers require to choose the best among all positively discriminated can…

Cited by 0SourceScholar