← Search

Haizhou Li

162 accepted papers

2026

ADAPTIVE PER-CHANNEL ENERGY NORMALIZATION FRONT-END FOR ROBUST AUDIO SIGNAL PROCESSING

ICASSP 2026poster

In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environment…

Cited by 0SourcePDFScholar
2026

AdaS: Adaptive Gradient Descent for Spiking Transformers

ICML 2026poster

Transformer-based Spiking Neural Networks (SNNs) combine Transformer performance with SNN energy efficiency through an event-driven self-attention mechanism. However, Spiking Transformers still lag behind their Artificial Neural Network (ANN) counterparts. Most existing studies address this issue th…

Cited by 0SourceScholar
2026

AdapAction: Adaptive Target Action Backdoor Attack against GUI Agents

CVPR 2026

Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) are increasingly vital for complex task automation. However, their capacity for self-driven decision-making introduces significant, yet underexplored, security risks, among which backdoor attacks pos

Cited by 0SourceScholar
2026

CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation

AAAI 2026technical

Theme detection is a fundamental task in user-centric dialogue systems, aiming to identify the latent topic of each utterance without relying on predefined schemas. Unlike intent induction, which operates within fixed label spaces, theme detection requires cross-dialogue consistency and alignment wi

Cited by 0SourcePDFScholar
2026

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

ICASSP 2026poster

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synthesis" methodology for training data construction. By generating source L2 speech…

Cited by 0SourcePDFScholar
2026

Direct Preference Optimization for Speech Autoregressive Diffusion Models

ICASSP 2026poster

Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressively generating continuous speech tokens with next-token diffusion, these models offer a promising alternative to next-to…

Cited by 0SourcePDFScholar
2026

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

ICLR 2026poster

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks t…

Cited by 0SourcecodeScholar
2026

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

IJCAI 2026

Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic--acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as

Cited by 0Scholar
2026

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

ICASSP 2026poster

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external…

Cited by 0SourcePDFScholar
2026

Neural Dynamics Self-Attention for Spiking Transformers

ICLR 2026poster

Integrating Spiking Neural Networks (SNNs) with Transformer architectures offers a promising pathway to balance energy efficiency and performance, particularly for edge vision applications. However, existing Spiking Transformers face two critical challenges: i) a substantial performance gap relative…

Cited by 0SourceScholar
2026

PAUL: Uncertainty-Guided Partition and Augmentation for Robust Cross-View Geo-Localization under Noisy Correspondence

CVPR 2026

Cross-view geo-localization is a critical task for UAV navigation, event detection, and aerial surveying, which establish correspondence between drone-captured and satellite imagery. Most existing approaches embed cross-view data into a joint feature space to maximize similarity between paired image

Cited by 0SourceScholar
2026

Positional Encoding for Spiking Transformers

ICML 2026poster

Spiking Neural Networks (SNNs) demonstrate superior energy efficiency over conventional Artificial Neural Networks (ANNs). Recent advances in Transformer-based SNNs have shown encouraging performance by seamlessly integrating spike-driven computation with Transformer architectures. Positional inform…

Cited by 0SourceScholar
2026

Robust Spiking Neural Networks Against Adversarial Attacks

ICLR 2026poster

Spiking Neural Networks (SNNs) represent a promising paradigm for energy-efficient neuromorphic computing due to their bio-plausible and spike-driven characteristics. However, the robustness of SNNs in complex adversarial environments remains significantly constrained. In this study, we theoretical…

Cited by 0SourceScholar
2026

SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation

ICML 2026spotlight

Spiking Neural Networks (SNNs) that leverage sparse binary spikes and temporal dynamics have emerged as energy-efficient alternatives to Artificial Neural Networks (ANNs). However, SNNs suffer from limited representational capacity due to the discrete nature of spikes. Existing solutions extending s…

Cited by 0SourceScholar
2026

SpikingLM: Towards Fully Spiking Language Model

ICML 2026poster

Spiking Neural Networks (SNNs) offer a promising avenue toward energy-efficient language modeling by replacing multiply-accumulate operations with sparse, event-driven computation. However, constructing fully spiking language models reveals two fundamental challenges: (1) gradient degradation from d…

Cited by 0SourceScholar
2026

TP-Spikformer: Token Pruned Spiking Transformer

ICLR 2026poster

Spiking neural networks (SNNs) offer an energy-efficient alternative to traditional neural networks due to their event-driven computing paradigm. However, recent advancements in spiking transformers have focused on improving accuracy with large-scale architectures, which require significant computat…

Cited by 0SourceScholar
2026

Towards Training-Free and Accurate ANN-to-SNN Conversion via Activation-Aware Redistribution

AAAI 2026technical

Conversion represents an effective approach for obtaining low-power models by transforming Artificial Neural Networks (ANNs) into event-driven Spiking Neural Networks (SNNs) without additional training. However, existing training-free conversion methods often incur substantial conversion errors. Her

Cited by 0SourcePDFScholar
2025

ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios

ICASSP 2025accepted

Decoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel mode…

Cited by 0SourceScholar
2025

Aligning Language Models Using Follow-up Likelihood as Reward Signal

AAAI 2025technical

In natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions,…

2025

Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation

EMNLP 2025

In the era of evaluating large language models (LLMs), data contamination has become an increasingly prominent concern. To address this risk, LLM benchmarking has evolved from a *static* to a *dynamic* paradigm. In this work, we conduct an in-depth analysis of existing *static* and *dynamic* benchma

2025

Binary Event-Driven Spiking Transformer

IJCAI 2025

Transformer-based Spiking Neural Networks (SNNs) introduce a novel event-driven self-attention paradigm that combines the high performance of Transformers with the energy efficiency of SNNs. However, the larger model size and increased computational demands of the Transformer structure limit their p

2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

ACL 2025finding

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding…

2025

ChatCRS: Incorporating External Knowledge and Goal Guidance for LLM-based Conversational Recommender Systems

NAACL 2025findings

This paper aims to efficiently enable large language models (LLMs) to use external knowledge and goal guidance in conversational recommender system (CRS) tasks. Advanced LLMs (e.g., ChatGPT) are limited in domain-specific CRS tasks for 1) generating grounded responses with recommendation-oriented kn…

Cited by 13SourcePDFScholar
2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge

NAACL 2025long

Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet often lack a robust methodology to dissect these phenomena comprehensively. Our work aims to bridge this gap by delving into the Food domain—a universally relevant yet culturally diverse aspect of hu…

2025

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

EMNLP 2025

The human-centered word association test (WAT) serves as a cognitive proxy, revealing sociocultural variations through culturally shared semantic expectations and implicit linguistic patterns shaped by lived experiences. We extend this test into an LLM-adaptive, free-relation task to assess the alig

2025

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

EMNLP 2025

Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions. To bridge this gap, we introduce Hanf

Cited by 0SourcePDFScholar
2025

Human Demonstrations are Generalizable Knowledge for Robots

IROS 2025

Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object

Cited by 11SourceScholar
2025

Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles

ACL 2025long

User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and…

2025

Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal Hypergraphs

NeurIPS 2025poster

Speech is a fundamental form of human communication, and speech perception constitutes the initial stage of language comprehension. Although brain-to-speech interface technologies have made significant progress in recent years, most existing studies focus on neural decoding during speech production.…

Cited by 0SourceScholar
2025

MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge

ICASSP 2025accepted

Multimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in da…

Cited by 0SourceScholar
2025

MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

ICASSP 2025accepted

In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same…

Cited by 0SourceScholar
2025

Multi-Level Speaker Representation for Target Speaker Extraction

ICASSP 2025accepted

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In th…

Cited by 0SourceScholar
2025

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

AAAI 2025technical

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual informat…

2025

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

EMNLP 2025

Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the targe

2025

Quantized Spike-driven Transformer

ICLR 2025poster

Spiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm. However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer struct…

2025

S$^2$NN: Sub-bit Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To furt…

Cited by 0SourceScholar
2025

Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion

ACL 2025long

This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or GPT-3.5, due to a predominant focus on mainstream languages (e.g., English and Ch…

2025

SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement

NeurIPS 2025poster

Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or s…

Cited by 0SourcecodeScholar
2025

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

AAAI 2025technical

The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about…

2025

Soundwave: Less is More for Speech-Text Alignment in LLMs

ACL 2025long

Existing end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth. We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency…

2025

Speech Separation for Low-Resource Languages

ICASSP 2025accepted

Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can lev…

Cited by 0SourceScholar
2025

Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models

NAACL 2025long

Data selection for fine-tuning large language models (LLMs) aims to choose a high-quality subset from existing datasets, allowing the trained model to outperform baselines trained on the full dataset. However, the expanding body of research lacks a clear, unified framework, and the variability in ex…

Cited by 5SourcePDFScholar
2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

ACL 2025long

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from multi-layer residual vector quantizer to single-layer quantizer are be…

2025

Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) demonstrate significant potential for energy-efficient neuromorphic computing through an event-driven paradigm. While training methods and computational models have greatly advanced, SNNs struggle to achieve competitive performance in visual long-sequence modeling task…

Cited by 0SourcecodeScholar
2024

A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators

AAAI 2024technical

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human eva…

2024

AceGPT, Localizing Large Language Models in Arabic

NAACL 2024long

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Significant concerns emerge when addressing cultural sensitivity and local values. T…

2024

Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark

COLING 2024main

Topic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and un…

2024

Alignment at Pre-training! Towards Native Alignment for Arabic LLMs

NeurIPS 2024poster

The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment…

2024

An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement

ICASSP 2024accepted

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encod…

Cited by 0SourceScholar
2024

Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition

IJCAI 2024poster

We introduce a novel multimodality synergistic knowledge distillation scheme tailored for efficient single-eye motion recognition tasks. This method allows a lightweight, unimodal student spiking neural network (SNN) to extract rich knowledge from an event-frame multimodal teacher network. The core…

Cited by 1SourcePDFScholar
2024

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech

ICASSP 2024accepted

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small…

Cited by 0SourceScholar
2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

EMNLP 2024finding

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. T…

2024

CMB: A Comprehensive Medical Benchmark in Chinese

NAACL 2024long

Large Language Models (LLMs) provide a possibility to make a great breakthrough in medicine. The establishment of a standardized medical benchmark becomes a fundamental cornerstone to measure progression. However, medical environments in different regions have their local characteristics, e.g., the…

2024

CrossTune: Black-Box Few-Shot Classification with Label Enhancement

COLING 2024main

Training or finetuning large-scale language models (LLMs) requires substantial computation resources, motivating recent efforts to explore parameter-efficient adaptation to downstream tasks. One approach is to treat these models as black boxes and use forward passes (Inference APIs) to interact with…

Cited by 3SourcePDFScholar
2024

DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language Models

EMNLP 2024main

Large language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting. However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers…

2024

Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling

AAAI 2024technical

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problem…

2024

Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise Ratio

ICASSP 2024accepted

Speaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Gra…

Cited by 0SourceScholar
2024

LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism

ICASSP 2024accepted

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization a…

Cited by 0SourceScholar
2024

Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading

NeurIPS 2024poster

Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy enviro…

2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 0SourceScholar
2024

LitE-SNN: Designing Lightweight and Efficient Spiking Neural Network through Spatial-Temporal Compressive Network Search and Joint Optimization

IJCAI 2024poster

Spiking Neural Networks (SNNs) mimic the information-processing mechanisms of the human brain and are highly energy-efficient, making them well-suited for low-power edge devices. However, the pursuit of accuracy in current studies leads to large, long-timestep SNNs, conflicting with the resource con…

Cited by 7SourcePDFScholar
2024

Prompt-Driven Target Speech Diarization

ICASSP 2024accepted

We introduce a novel task named ‘target speech diarization’, which seeks to determine ‘when target event occurred’ within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (PTSD), that works with diverse prompts that specify the target speech events of i…

Cited by 0SourceScholar
2024

Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition

AAAI 2024technical

Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a vid…

Cited by 11SourcePDFScholar
2024

Robust Decoding of the Auditory Attention from EEG Recordings Through Graph Convolutional Networks

ICASSP 2024accepted

Auditory attention decoding (AAD) with electroencephalography (EEG) holds great promise in brain-computer interface (BCI). Despite much progress, it remains a research topic on how to effectively evaluate the performance of EEG-based AAD algorithms under an appropriate setting that reflects the use…

Cited by 0SourceScholar
2024

SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

NeurIPS 2024poster

Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction. Chat-Oriented Large Language Models (LLMs), known fo…

2024

SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

ICASSP 2024accepted

Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achi…

Cited by 0SourceScholar
2024

Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural Networks

ICASSP 2024accepted

Brain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due to the lack of an effective auditory front-end. To address this limitation, we introduce Spiking-LEAF, a learnable audit…

Cited by 0SourceScholar
2024

TC-LIF: A Two-Compartment Spiking Neuron Model for Long-Term Sequential Modelling

AAAI 2024technical

The identification of sensory cues associated with potential opportunities and dangers is frequently complicated by unrelated events that separate useful cues by long delays. As a result, it remains a challenging task for state-of-the-art spiking neural networks (SNNs) to establish long-term tempora…

2024

TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models

EMNLP 2024finding

Mainstream approaches to aligning large language models (LLMs) heavily rely on human preference data, particularly when models require periodic updates. The standard process for iterative alignment of LLMs involves collecting new human feedback for each update. However, the data collection process i…

2024

UNO-DST: Leveraging Unlabelled Data in Zero-Shot Dialogue State Tracking

NAACL 2024findings

Previous zero-shot dialogue state tracking (DST) methods only apply transfer learning, but ignore unlabelled data in the target domain.We transform zero-shot DST into few-shot DST by utilising such unlabelled data via joint and self-training methods. Our method incorporates auxiliary tasks that gene…

2024

Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical Study

COLING 2024main

Large language models, like ChatGPT, have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored, where it requires higher level capabilities of understanding and reasoning. In this paper, we aim to systematically…

2024

Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models

ACL 2024findings

The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness of these evaluators against adversarial perturbations remains largely under-expl…

2024

VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions

NeurIPS 2024poster

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, c…

Cited by 5SourcePDFScholar
2023

Disentangling Voice and Content with Self-Supervision for Speaker Recognition

NeurIPS 2023poster

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with t…

Cited by 51SourcePDFScholar
2023

Dynamic Transformers Provide a False Sense of Efficiency

ACL 2023long

Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference. Multi-exit is a mainstream approach to address this issue by making a trade-off between efficiency and accuracy, where the saving of computation comes…

2023

Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing Modalities

ICASSP 2023accepted

Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing data across modalities, the inherent difference between heterog…

Cited by 0SourceScholar
2023

How Well Do Text Embedding Models Understand Syntax?

EMNLP 2023long findings

Text embedding models have significantly contributed to advancements in natural language processing by adeptly capturing semantic properties of textual data. However, the ability of these models to generalize across a wide range of syntactic contexts remains under-explored. In this paper, we first d…

Cited by 0SourcecodeScholar
2023

HuatuoGPT, Towards Taming Language Model to Be a Doctor

EMNLP 2023long findings

In this paper, we present HuatuoGPT, a Large Language Model (LLM) for medical consultation. The core recipe of HuatuoGPT is to leverage both distilled data from **ChatGPT** and real-world data from **doctors** in the supervised fine-tuning stage. This is not only because purely using **ChatGPT**-di…

Cited by 0SourcecodeScholar
2023

ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting

ICASSP 2023accepted

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video cl…

Cited by 0SourceScholar
2023

Minimizing the Accumulated Trajectory Error To Improve Dataset Distillation

CVPR 2023poster

Model-based deep learning has achieved astounding successes due in part to the availability of large-scale real-world data. However, processing such massive amounts of data comes at a considerable cost in terms of computations, storage, training and the search for good neural architectures. Dataset…

2023

Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response

ICASSP 2023accepted

This work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefi…

Cited by 0SourceScholar
2023

Ripple Sparse Self-Attention for Monaural Speech Enhancement

ICASSP 2023accepted

The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the…

Cited by 10SourceScholar
2023

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

CVPR 2023poster

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip move…

2023

Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and Text

ICASSP 2023accepted

Self-supervised pre-training has been successful in both text and speech processing. Speech and text offer different but complementary information. The question is whether we are able to perform a speech-text joint pre-training on unpaired speech and text. In this paper, we take the idea of self-sup…

Cited by 0SourceScholar
2023

xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark

EMNLP 2023long findings

Recent advancements in reference-free learned metrics for open-domain dialogue evaluation have been driven by the progress in pre-trained language models and the availability of dialogue data with high-quality human annotations. However, current studies predominantly concentrate on English dialogues…

Cited by 0SourcecodeScholar
2022

A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal Coding

ICASSP 2022accepted

Bio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in spars…

Cited by 0SourceScholar
2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

Analyzing and Evaluating Faithfulness in Dialogue Summarization

EMNLP 2022main

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text summarization. However, there is a lack of systematic study…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages

ICASSP 2022accepted

Target language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, si…

Cited by 0SourceScholar
2022

FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation

EMNLP 2022main

Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality di…

2022

Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning Framework

EMNLP 2022main

Most sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals. Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks. In thi…

2022

Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic Music

ICASSP 2022accepted

Lyrics transcription of polyphonic music is challenging not only because the singing vocals are corrupted by the background music, but also because the background music and the singing style vary across music genres, such as pop, metal, and hip hop, which affects lyrics intelligibility of the song i…

Cited by 0SourceScholar
2022

L-SpEx: Localized Target Speaker Extraction

ICASSP 2022accepted

Speaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s…

Cited by 0SourceScholar
2022

M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database

ACL 2022long

The emotional state of a speaker can be influenced by many different factors in dialogues, such as dialogue scene, dialogue topic, and interlocutor stimulus. The currently available data resources to support such multimodal affective analysis in dialogues are however limited in scale and diversity.…

2022

MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation

AAAI 2022technical

Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure the quality of such conversational agents, a dialogue evaluator is expected to conduct assessment across domains as well…

2022

MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short Utterances

ICASSP 2022accepted

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems…

Cited by 0SourceScholar
2022

Memobert: Pre-Training Model with Prompt-Based Learning for Multimodal Emotion Recognition

ICASSP 2022accepted

Multimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity. In this paper, we propose a multimodal pre-training model MEmoBERT for multimodal emotion recognition, which learns multimodal joint…

Cited by 0SourceScholar
2022

Self-Supervised Speaker Recognition with Loss-Gated Learning

ICASSP 2022accepted

In self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to mod…

Cited by 0SourceScholar
2022

Time-Frequency Attention for Monaural Speech Enhancement

ICASSP 2022accepted

Most studies on speech enhancement generally don’t explicitly consider the energy distribution of speech in time-frequency (T-F) representation, which is important for accurate prediction of mask or spectra. In this paper, we present a simple yet effective T-F attention (TFA) module, where a…

Cited by 35SourceScholar
2022

Training Spiking Neural Networks with Local Tandem Learning

NeurIPS 2022accept

Spiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a g…

2022

Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over

ICASSP 2022accepted

In this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but also perfect lip-speech synchronization. A natural solution to…

Cited by 0SourceScholar
2021

Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural Networks

ICML 2021spotlight

Gradient staleness is a major side effect in decoupled learning when training convolutional neural networks asynchronously. Existing methods that ignore this effect might result in reduced generalization and even divergence. In this paper, we propose an accumulated decoupled learning (ADL), which in…

2021

Bootstrapped Unsupervised Sentence Representation Learning

ACL 2021long

As high-quality labeled data is scarce, unsupervised sentence representation learning has attracted much attention. In this paper, we propose a new framework with a two-branch Siamese Network which maximizes the similarity between two augmented views of each sentence. Specifically, given one augment…

2021

Data Augmentation with Signal Companding for Detection of Logical Access Attacks

ICASSP 2021accepted

The recent advances in voice conversion (VC) and text-to-speech (TTS) make it possible to produce natural sounding speech that poses threat to automatic speaker verification (ASV) systems. To this end, research on spoofing countermeasures has gained attention to protect ASV systems from such attacks…

Cited by 0SourceScholar
2021

DynaEval: Unifying Turn and Dialogue Level Evaluation

ACL 2021long

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring such dynamics. To this end, we propose DynaEval, a unified…

2021

GCC-PHAT with Speech-oriented Attention for Robotic Sound Source Localization

ICRA 2021poster

Robotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments du…

Cited by 19SourceScholar
2021

Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training

ICASSP 2021accepted

Neural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the…

Cited by 0SourceScholar
2021

Leveraging Acoustic and Linguistic Embeddings from Pretrained Speech and Language Models for Intent Classification

ICASSP 2021accepted

Intent classification is a task in spoken language understanding. An intent classification system is usually implemented as a pipeline process, with a speech recognition module followed by text processing that classifies the intents. There are also studies of end-to-end system that take acoustic fea…

Cited by 24SourceScholar
2021

Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals

ICASSP 2021accepted

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted s…

Cited by 0SourceScholar
2021

Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism

ICASSP 2021accepted

Most of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker local…

Cited by 0SourceScholar
2021

Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition

ICASSP 2021accepted

Convolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spe…

Cited by 0SourceScholar
2021

Revisiting Self-training for Few-shot Learning of Language Model

EMNLP 2021main

As unlabeled data carry rich task-relevant information, they are proven useful for few-shot learning of language model. The question is how to effectively make use of such data. In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-…

2021

Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech Dataset

ICASSP 2021accepted

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network conditioned on discrete representation, such as one-hot emot…

Cited by 0SourceScholar
2021

The Multi-Speaker Multi-Style Voice Cloning Challenge 2021

ICASSP 2021accepted

The Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited d…

Cited by 0SourceScholar
2020

Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?

ICASSP 2020accepted

Automatic lyrics alignment and transcription in polyphonic music are challenging tasks because the singing vocals are corrupted by the background music. In this work, we propose to learn music genre-specific characteristics to train polyphonic acoustic models. We first compare several automatic spee…

Cited by 0SourceScholar
2020

End-to-End Code-Switching TTS with Cross-Lingual Language Model

ICASSP 2020accepted

Code-switching text-to-speech (TTS) aims to enable a system to speak two languages with a single voice and in the same utterance. In this paper, we propose to incorporate cross-lingual word embedding into an end-to-end TTS system, to improve the voice rendering. The cross-lingual word embedding, gen…

Cited by 0SourceScholar
2020

Independent Language Modeling Architecture for End-To-End ASR

ICASSP 2020accepted

The attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model…

Cited by 0SourceScholar
2020

On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech Study

ICASSP 2020accepted

Characterizing speakers under stressed condition is a challenge because speakers deviate from the normal speech production process. Whispered speech is one among them that is produced by abducting the vocal folds to pass the air out of mouth. During this process, the airflow thus passed is influence…

Cited by 0SourceScholar
2020

Teacher-Student Training For Robust Tacotron-Based TTS

ICASSP 2020accepted

While neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that resul…

Cited by 0SourceScholar
2020

Time-Domain Neural Network Approach for Speech Bandwidth Extension

ICASSP 2020accepted

In this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of informat…

Cited by 0SourceScholar
2019

Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection

ICASSP 2019accepted

The security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. T…

Cited by 0SourceScholar
2019

Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic Models

ICASSP 2019accepted

Lyrics-to-audio alignment is to automatically align the lyrical words with the mixed singing audio (singing voice+musical accompaniment). Such alignment can be achieved with an automatic speech recognition (ASR) system. We propose to adapt the acoustic model of a speech recognizer towards solo singi…

Cited by 0SourceScholar
2019

Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average Modeling

ICASSP 2019accepted

This paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In partic…

Cited by 0SourceScholar
2019

Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss

ICASSP 2019accepted

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and cons…

Cited by 0SourceScholar
2018

End-to-End Hierarchical Language Identification System

ICASSP 2018accepted

Recently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure t…

Cited by 0SourceScholar
2018

On the Importance of Analytic Phase of Speech Signals in Spoken Language Recognition

ICASSP 2018accepted

In this paper, we study the role of long-time analytic phase of speech signals in spoken language recognition (SLR) and employ a set of features termed as instantaneous frequency cepstral coefficients (IFCC). We extract IFCC from long-time analytic phase, in an effort to capture long range acoustic…

Cited by 0SourceScholar
2018

Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTM

ICASSP 2018accepted

Utterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. Howev…

Cited by 0SourceScholar
2018

Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition

ICASSP 2018accepted

The i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degr…

Cited by 0SourceScholar
2017

Adaptation of PLDA for multi-source text-independent speaker verification

ICASSP 2017accepted

Probabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, th…

Cited by 0SourceScholar
2017

On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition

ICASSP 2017accepted

Acoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estim…

Cited by 0SourceScholar
2017

Pairwise learning using multi-lingual bottleneck features for low-resource query-by-example spoken term detection

ICASSP 2017accepted

We propose to use a feature representation obtained by pairwise learning in a low-resource language for query-by-example spoken term detection (QbE-STD). We assume that word pairs identified by humans are available in the low-resource target language. The word pairs are parameterized by a multi-ling…

Cited by 0SourceScholar
2016

A hierarchical framework for language identification

ICASSP 2016accepted

Most current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable…

Cited by 0SourceScholar
2016

An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources

ICASSP 2016accepted

This paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector c…

Cited by 0SourceScholar
2016

Approximate search of audio queries by using DTW with phone time boundary and data augmentation

ICASSP 2016accepted

Dynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DT…

Cited by 0SourceScholar
2016

Combining multiple kernel models for automatic intelligibility detection of pathological speech

ICASSP 2016accepted

Automatic detection of pathological voice is a challenging task in speech processing. Appropriate acoustic cues of voice can be used to differentiate between normal voices and pathological voices. We propose a method to represent each speech utterance using three types of speech signal representatio…

Cited by 4SourceScholar
2016

Content-aware local variability vector for speaker verification with short utterance

ICASSP 2016accepted

I-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session var…

Cited by 0SourceScholar
2016

Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search

ICASSP 2016accepted

In this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the…

Cited by 0SourceScholar
2016

Exemplar-based sparse representation of timbre and prosody for voice conversion

ICASSP 2016accepted

Voice conversion (VC) aims to make one speaker (source) to sound like spoken by another speaker (target) without changing the language content. Most of the state-of-the-art voice conversion systems focus only on timbre conversion. However, the speaker identity is characterized by the source-related…

Cited by 0SourceScholar
2016

Exemplar-inspired strategies for low-resource spoken keyword search in Swahili

ICASSP 2016accepted

We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational…

Cited by 0SourceScholar
2016

Keyword search using query expansion for graph-based rescoring of hypothesized detections

ICASSP 2016accepted

In this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speec…

Cited by 0SourceScholar
2016

Spoofing detection from a feature representation perspective

ICASSP 2016accepted

Spoofing detection, which discriminates the spoofed speech from the natural speech, has gained much attention recently. Low-dimensional features that are used in speaker recognition/verification are also used in spoofing detection. Unfortunately, they don't capture sufficient information required fo…

Cited by 0SourceScholar
2015

A learning-based approach to direction of arrival estimation in noisy and reverberant environments

ICASSP 2015accepted

This paper presents a learning-based approach to the task of direction of arrival estimation (DOA) from microphone array input. Traditional signal processing methods such as the classic least square (LS) method rely on strong assumptions on signal models and accurate estimations of time delay of arr…

Cited by 0SourceScholar
2015

Channel adaptation of plda for text-independent speaker verification

ICASSP 2015accepted

Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log…

Cited by 0SourceScholar
2015

Combining robust spike coding with spiking neural networks for sound event classification

ICASSP 2015accepted

This paper proposes a novel biologically inspired method for sound event classification which combines spike coding with a spiking neural network (SNN). Our spike coding extracts keypoints that represent the local maxima components of the sound spectrogram, and are encoded based on their local time-…

Cited by 0SourceScholar
2015

Language independent query-by-example spoken term detection using N-best phone sequences and partial matching

ICASSP 2015accepted

In this paper, we propose a partial sequence matching based symbolic search (SS) method for the task of language independent query-by-example spoken term detection. One main drawback of conventional SS approach is the high miss rate for long queries. This is due to high variations in symbol represen…

Cited by 0SourceScholar
2015

Low-resource keyword search strategies for tamil

ICASSP 2015accepted

We propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological…

Cited by 0SourceScholar
2015

Source-specific informative prior for i-vector extraction

ICASSP 2015accepted

An i-vector is a low-dimensional fixed-length representation of a variable-length speech utterance, and is defined as the posterior mean of a latent variable conditioned on the observed feature sequence of an utterance. The assumption is that the prior for the latent variable is non-informative, sin…

Cited by 0SourceScholar
2015

Tokenizing fundamental frequency variation for Mandarin tone error detection

ICASSP 2015accepted

Tone error is commonly observed in tonal language acquisition. Correct tone production is especially challenging for native speakers of non-tonal languages. In this paper, we exploit the fundamental frequency variation (FFV) feature for Mandarin tone error detection. We propose to use FFV through tw…

Cited by 0SourceScholar