← Search

Kuan-Yu Chen

17 accepted papers

2026

Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks

ICLR 2026poster

Text embeddings enable numerous NLP applications but face severe privacy risks from embedding inversion attacks, which can expose sensitive attributes or reconstruct raw text. Existing differential privacy defenses assume uniform sensitivity across embedding dimensions, leading to excessive noise an…

Cited by 0SourceScholar
2026

Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems

ICASSP 2026oral

Instruction-guided text-to-speech (ITTS) enables users to control speech generation through natural language prompts, offering a more intuitive interface than traditional TTS. However, the alignment between user style instructions and listener perception remains largely unexplored. This work first p…

Cited by 1SourcePDFScholar
2026

Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement

ICLR 2026poster

Existing large language model (LLM)-based embeddings typically adopt an encoder-only paradigm, treating LLMs as static feature extractors and overlooking their core gener- ative strengths. We introduce GIRCSE (Generative Iterative Refinement for Contrastive Sentence Embeddings), a novel framework th…

Cited by 0SourceScholar
2025

Creativity in LLM-based Multi-Agent Systems: A Survey

EMNLP 2025

Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of creativity , including how novel outputs

2024

DOFEN: Deep Oblivious Forest ENsemble

NeurIPS 2024poster

Deep Neural Networks (DNNs) have revolutionized artificial intelligence, achieving impressive results on diverse data types, including images, videos, and texts. However, DNNs still lag behind Gradient Boosting Decision Trees (GBDT) on tabular data, a format extensively utilized across various domai…

2023

Trompt: Towards a Better Deep Neural Network for Tabular Data

ICML 2023poster

Tabular data is arguably one of the most commonly used data structures in various practical domains, including finance, healthcare and e-commerce. The inherent heterogeneity allows tabular data to store rich information. However, based on a recently published tabular benchmark, we can see deep neura…

Cited by 33SourcePDFScholar
2021

Speech Recognition by Simply Fine-Tuning Bert

ICASSP 2021accepted

We propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate rich contextual representations. Our assumption is that given a history context sequence, a powerful LM can narrow the ran…

Cited by 0SourceScholar
2018

Essence Vector-Based Query Modeling for Spoken Document Retrieval

ICASSP 2018accepted

Spoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods an…

Cited by 0SourceScholar
2018

Scalable Sentiment for Sequence-to-Sequence Chatbot Response with Performance Analysis

ICASSP 2018accepted

Conventional seq2seq chatbot models only try to find the sentences with the highest probabilities conditioned on the input sequences, without considering the sentiment of the output sentences. Some research works trying to modify the sentiment of the output sequences were reported. In this paper, we…

Cited by 35SourceScholar
2017

A locality-preserving essence vector modeling framework for spoken document retrieval

ICASSP 2017accepted

Because unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machi…

Cited by 0SourceScholar
2017

Leveraging manifold learning for extractive broadcast news summarization

ICASSP 2017accepted

Extractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech…

Cited by 0SourceScholar
2016

Daily activity recognition using the informative features from skeletal and depth data

ICRA 2016

In this paper, we present an efficient framework for human activity recognition in daily environment. We use depth information mainly for privacy protection, and then focus on the motion analysis of informative body parts, since most activities are much associated with these particular parts, e.g.,

Cited by 8SourceScholar
2016

Improved spoken document summarization with coverage modeling techniques

ICASSP 2016accepted

Extractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods desi…

Cited by 0SourceScholar