← Search

Zhongqin Wu

10 accepted papers

2022

A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction

ICASSP 2022accepted

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based…

Cited by 0SourceScholar
2022

All-in-One Image Restoration for Unknown Corruption

CVPR 2022poster

In this paper, we study a challenging problem in image restoration, namely, how to develop an all-in-one method that could recover images from a variety of unknown corruption types and levels. To this end, we propose an All-in-one Image Restoration Network (AirNet) consisting of two neural modules,…

Cited by 351PDFcodeScholar
2022

Retrieval-Based Spatially Adaptive Normalization for Semantic Image Synthesis

CVPR 2022poster

Semantic image synthesis is a challenging task with many practical applications. Albeit remarkable progress has been made in semantic image synthesis with spatially-adaptive normalization and existing methods normalize the feature activations under the coarse-level guidance (e.g., semantic class). H…

Cited by 33PDFcodeScholar
2022

Syntax-Aware Network for Handwritten Mathematical Expression Recognition

CVPR 2022poster

Handwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods adhere to the paradigm that the prediction is made "from one c…

Cited by 98PDFScholar
2022

Time-Domain Audio-Visual Speech Separation on Low Quality Videos

ICASSP 2022accepted

Incorporating visual information is a promising approach to improve the performance of speech separation. Many related works have been conducted and provide inspiring results. However, low quality videos appear commonly in real scenarios, which may significantly degrade the performance of normal aud…

Cited by 0SourceScholar
2021

CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations

EMNLP 2021main

Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms. However, these models are facing challenges of overfitting with limited labels and low model generalization abilities. In this paper, we present a Cross-modal Transformer for Audio-and-L…

2021

FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation

CVPR 2021poster

Image caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whethe…

Cited by 41PDFcodeScholar
2021

Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations

EMNLP 2021main

There is an increasing interest in the use of mathematical word problem (MWP) generation in educational assessment. Different from standard natural question generation, MWP generation needs to maintain the underlying mathematical operations between quantities and variables, while at the same time en…

2021

Orthogonal Jacobian Regularization for Unsupervised Disentanglement in Image Generation

ICCV 2021poster

Unsupervised disentanglement learning is a crucial issue for understanding and exploiting deep generative models. Recently, SeFa tries to find latent disentangled directions by performing SVD on the first projection of a pre-trained GAN. However, it is only applied to the first layer and works in a…

Cited by 72PDFcodeScholar
2020

Personalized Multimodal Feedback Generation in Education

COLING 2020main

The automatic feedback of school assignments is an important application of AI in education. In this work, we focus on the task of personalized multimodal feedback generation, which aims to generate personalized feedback for teachers to evaluate students’ assignments involving multimodal inputs such…