← Search

Feng-Ju Chang

9 accepted papers

2025

Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework

ICCV 2025poster

We present Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM), a framework that advances vision-language matching and retrieval by leveraging a large language model (LLM) backbone. While concurrent LLM-based approaches like VLM2VEC, MM-Embed, NV-Embed, and MM-GEM have demonstrated impressive c…

2024

A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis

ICLR 2024poster

We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an im…

2023

Dialog Act Guided Contextual Adapter for Personalized Speech Recognition

ICASSP 2023accepted

Personalization in multi-turn dialogs has been a long standing challenge for end-to-end automatic speech recognition (E2E ASR) models. Recent work on contextual adapters has tackled rare word recognition using user catalogs. This adaptation, however, does not incorporate an important cue, the dialog…

Cited by 0SourceScholar
2023

Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition

ICASSP 2023accepted

We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its att…

Cited by 6SourceScholar
2023

Gated Contextual Adapters For Selective Contextual Biasing In Neural Transducers

ICASSP 2023accepted

Neural contextual biasing for end-to-end neural ASR transducers has shown significant improvements in the recognition of named entities, such as contact names or device names. However, it comes with the cost of increased compute, as the biasing layers (which are usually based on cross-attention) add…

Cited by 12SourceScholar
2022

Contextual Adapters for Personalized Speech Recognition in Neural Transducers

ICASSP 2022accepted

Personal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is with shallow fusion methods at inference time. However, due to their dependence on external language models and the dete…

Cited by 0SourceScholar
2022

Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

ICASSP 2022accepted

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally cascaded pipelines. Existing E2E SLU models usually follow a two-stage configuration where an Automatic Speech Recognition (…

Cited by 0SourceScholar
2021

End-to-End Multi-Channel Transformer for Speech Recognition

ICASSP 2021accepted

Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the spectral and spatial information collected from different mic…

Cited by 0SourceScholar
2021

Sparsification via Compressed Sensing for Automatic Speech Recognition

ICASSP 2021accepted

In order to achieve high accuracy for machine learning (ML) applications, it is essential to employ models with a large number of parameters. Certain applications, such as Automatic Speech Recognition (ASR), however, require real-time interactions with users, hence compelling the model to have as lo…

Cited by 0SourceScholar