← Search

Sakriani Sakti

17 accepted papers

2026

WAVENEXT 2: CONVNEXT-BASED FAST NEURAL VOCODERS WITH RESIDUAL DENOISING AND SUB-MODELING FOR GAN AND DIFFUSION MODELS

ICASSP 2026poster

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited performance in multi-speaker settings. Moreover, diffusion models, de…

Cited by 0SourcePDFScholar
2025

Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes

ICASSP 2025accepted

Most recent Acoustic Word Embedding (AWE) systems utilize an autoencoder-like approach to compress speech features of arbitrary shapes into fixed-size numerical vectors and then reconstructing it, thereby capturing essential patterns in the data. Unfortunately, AWE models have commonly relied on sup…

Cited by 0SourceScholar
2025

From Pixels to Voice: A Simple and Efficient End-to-End Spoken Image Description Approach via Vision Codec Language Models

ICASSP 2025accepted

Neural audio codecs provide a powerful tool for compressing audio signals into discrete codec representations. This compact discrete representation has made it possible to successfully apply a natural language processing (NLP) model to various audio and speech processing tasks, including text-to-spe…

Cited by 0SourceScholar
2025

Toward Visual Pronunciation Learning: A Speech-to-Articulatory Animation Pipeline Leveraging wav2vec 2.0 and rtMRI Landmarks

ICASSP 2025accepted

Most computer-assisted pronunciation training (CAPT) systems for second language (L2) learners focus on detecting mispronunciation based on predefined phonemes and assigning pronunciation scores. However, these systems often lack visual feedback or detailed corrective guidance, limiting learners’ op…

Cited by 0SourceScholar
2024

Refining rtMRI Landmark-Based Vocal Tract Contour Labels with FCN-Based Smoothing and Point-to-Curve Projection

COLING 2024main

Advanced real-time Magnetic Resonance Imaging (rtMRI) enables researchers to study dynamic articulatory movements during speech production with high temporal resolution. However, accurately outlining articulator contours in high-frame-rate rtMRI presents challenges due to data scalability and image…

Cited by 1SourcePDFScholar
2023

An Isotropy Analysis for Self-Supervised Acoustic Unit Embeddings on the Zero Resource Speech Challenge 2021 Framework

ICASSP 2023accepted

In recent years, self-supervised representation learning has gained much attention for its proven advantages in many downstream tasks. Consequently, various self-supervised representation learning methods have been developed. However, few studies have investigated the resulting embedding space or an…

Cited by 0SourceScholar
2023

NusaCrowd: Open Source Initiative for Indonesian NLP Resources

ACL 2023findings

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the dataset…

2023

Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise Condition

ICASSP 2023accepted

A common approach for text-to-speech (TTS) in noisy conditions is offline fine-tuning, which is generally utilized on static noises and predefined conditions. We recently proposed a self-adaptive TTS in machine speech chain inference that enables TTS to control its voices in statically and dynamical…

Cited by 0SourceScholar
2023

Speech Recognition and Meaning Interpretation: Towards Disambiguation of Structurally Ambiguous Spoken Utterances in Indonesian

EMNLP 2023long main

Despite being the world's fourth-most populous country, the development of spoken language technologies in Indonesia still needs improvement. Most automatic speech recognition (ASR) systems that have been developed are still limited to transcribing the exact word-by-word, which, in many cases, consi…

Cited by 0SourcecodeScholar
2020

Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinate

ICASSP 2020accepted

3D panoramic multi-person localization and tracking are prominent in many applications, however, conventional methods using LiDAR equipment could be economically expensive and also computationally inefficient due to the processing of point cloud data. In this work, we propose an effective and effici…

Cited by 0SourceScholar
2019

Cross-lingual Speech-based Tobi Label Generation Using Bidirectional Lstm

ICASSP 2019accepted

In this paper we investigate the automatic generation of ToBI-style prosody labels. The work is motivated by the idea of using prosodic information to facilitate the automatic lexicon discovery for unseen and under-resourced languages for which sufficient training data is not available. Specifically…

Cited by 0SourceScholar
2019

End-to-end Feedback Loss in Speech Chain Framework via Straight-through Estimator

ICASSP 2019accepted

The speech chain mechanism integrates automatic speech recognition (ASR) and text-to-speech synthesis (TTS) modules into a single cycle during training. In our previous work, we applied a speech chain mechanism as a semi-supervised learning. It provides the ability for ASR and TTS to assist each oth…

Cited by 0SourceScholar
2019

Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor Decomposition

ICASSP 2019accepted

Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG…

Cited by 0SourceScholar
2018

Graph Regularized Tensor Factorization for Single-Trial EEG Analysis

ICASSP 2018accepted

This study proposes a tensor factorization algorithm for electroencephalographies (EEGs) that incorporates the geometric structure of the electrode location. The purpose is removing noise caused by EEG activities which are irrelevant to stimuli presented to a subject from single-trial event-related…

Cited by 0SourceScholar
2015

Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASR

ICASSP 2015accepted

This paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks…

Cited by 22SourceScholar
2015

EEG signal enhancement using multi-channel wiener filter with a spatial correlation prior

ICASSP 2015accepted

Event-related potentials (ERPs) of electroencephalogram (EEG) are often used as features for brain machine interfaces or for analysis of brain activities. However, as EEG signals easily suffer from various artifacts, ERPs are often collapsed and hard to observe. There are several attempts at using m…

Cited by 0SourceScholar