← Search

Ming Fang

10 accepted papers

2025

A Teacher Action Quality Assessment Method Based on Label Constraint Strategy

ICASSP 2025accepted

Teacher’s actions significantly impact teaching effectiveness. Assessing these actions can help teachers to identify shortcomings and improve skills. However, the lack of datasets for assessing teacher action quality has hindered progress in this field. Therefore, this paper first constructs the Tea…

Cited by 0SourceScholar
2025

EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed

ICASSP 2025accepted

Non-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR A…

Cited by 0SourceScholar
2025

Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss

ICASSP 2025accepted

Contextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the…

Cited by 0SourceScholar
2025

LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream Generator

ICASSP 2025accepted

Recently, the field of Text-to-speech synthesis has been predominantly characterized by end-to-end models, with the quality of speech generated by these models becoming increasingly comparable to that of human speech. In this work, we propose a Lightweight and Efficient Text-to-speech model, a fast…

Cited by 0SourceScholar
2025

Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASR

ICASSP 2025accepted

Contextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual n…

Cited by 0SourceScholar
2024

A Teacher Classroom Dress Assessment Method Based on a New Assessment Dataset

IJCAI 2024poster

Proper attire is a professional requirement for teachers and teachers' dress influence students' perceptions of teacher quality. Therefore, evaluating teacher attire can better regulate and improve the teacher’s dress. However, the lack of a dataset on teacher attire hinders the development of this…

2024

Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix Reconstruction

ICASSP 2024accepted

In automatic speech recognition (ASR) task, the output sequence should correspond to a linear transcription of the input sequence. Lots of works have been done to learn the monotonic alignment in end-to-end (E2E) ASR model, but their methods mainly focus on streaming propose and usually result in a…

Cited by 0SourceScholar
2024

Which is the Better Teacher Action? A New Ranking Model and Dataset

ICASSP 2024accepted

Teachers as leaders of classroom teaching, can enhance students’ learning interest by effectively using body language. Consequently, the quality of teachers’ actions is one of the critical factors influencing the teaching effect. Teachers can find their shortcomings and improve their teaching skills…

Cited by 0SourceScholar
2022

Analyzing the Intensity of Complaints on Social Media

NAACL 2022findings

Complaining is a speech act that expresses a negative inconsistency between reality and human’s expectations. While prior studies mostly focus on identifying the existence or the type of complaints, in this work, we present the first study in computational linguistics of measuring the intensity of c…

2021

SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech Recognition

ICASSP 2021accepted

Inspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximi…

Cited by 0SourceScholar