← Search

Brian Mak

16 accepted papers

2025

End-to-End Optimization for Multimodal Retrieval-Augmented Generation via Reward Backpropagation

EMNLP 2025

Multimodal Retrieval-Augmented Generation (MM-RAG) has emerged as a promising approach for enhancing the reliability and factuality of large vision-language models (LVLMs). While end-to-end loss backpropagation is infeasible due to non-differentiable operations during the forward process, current me

2024

A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News

COLING 2024main

This paper introduces TVB-HKSL-News, a new Hong Kong sign language (HKSL) dataset collected from a TV news program over a period of 7 months. The dataset is collected to enrich resources for HKSL and support research in large-vocabulary continuous sign language recognition (SLR) and translation (SLT…

Cited by 4SourcePDFScholar
2024

A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars

ECCV 2024oral

"The objective of this paper is to develop a functional system for translating spoken languages into sign languages, referred to as Spoken2Sign translation. The Spoken2Sign task is orthogonal and complementary to traditional sign language to spoken language (Sign2Spoken) translation. To enable Spoke…

2022

Two-Stream Network for Sign Language Recognition and Translation

NeurIPS 2022accept

Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substanti…

2021

A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection

ICASSP 2021accepted

With the global population ageing rapidly, Alzheimer's disease (AD) is particularly prominent in older adults, which has an insidious onset followed by gradual, irreversible deterioration in cognitive domains (memory, communication, etc). Thus the detection of Alzheimer's disease is crucial for time…

Cited by 0SourceScholar
2020

Stochastic Fine-grained Labeling of Multi-state Sign Glosses for Continuous Sign Language Recognition

ECCV 2020poster

In this paper, we propose novel stochastic modeling of various components of a continuous sign language recognition (CSLR) system that is based on the transformer encoder and connectionist temporal classification (CTC). Most importantly, We model each sign gloss with multiple states, and the number…

2018

learning Effective Factorized Hidden Layer Bases Using Student-Teacher Training for LSTM Acoustic Model Adaptation

ICASSP 2018accepted

Factorized Hidden Layer (FHL) has been proposed for the adaptation of deep neural network (DNN) and Long Short-Term Memory (LSTM) based acoustic models (AMs). In FHL, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD…

Cited by 0SourceScholar
2017

An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptation

ICASSP 2017accepted

Subspace methods are used for deep neural network (DNN)-based acoustic model adaptation. These methods first construct a subspace and then perform the speaker adaptation as a point in the subspace. This paper aims to investigate the effectiveness of subspace methods for robust unsupervised adaptatio…

Cited by 0SourceScholar
2017

Speeding up softmax computations in DNN-based large vocabulary speech recognition by senone weight vector selection

ICASSP 2017accepted

Deep neural network has obtained significant accuracy improvement in many large vocabulary continuous speech recognition (LVCSR) tasks. Recently, it was shown that even better performance can be obtained by modeling a larger number of more discriminative senones. However, as the neural network becom…

Cited by 0SourceScholar