← Search

Ke Gong

10 accepted papers

2024

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

AAAI 2024technical

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, of…

2021

Adversarial Meta Sampling for Multilingual Low-Resource Speech Recognition

AAAI 2021technical

Low-resource automatic speech recognition (ASR) is challenging, as the low-resource target language data cannot well train an ASR model. To solve this issue, meta-learning formulates ASR for each source language into many small ASR tasks and meta-learns a model initialization on all tasks from diffe…

Cited by 35SourcePDFScholar
2021

Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition

EMNLP 2021finding

Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-resource speech recognition. Existing approaches simply cascade pre-trained acoustic and language models to learn the tra…

2020

Bidirectional Graph Reasoning Network for Panoptic Segmentation

CVPR 2020poster

Recent researches on panoptic segmentation resort to a single end-to-end network to combine the tasks of instance segmentation and semantic segmentation. However, prior models only unified the two related tasks at the architectural level via a multi-branch scheme or revealed the underlying correlati…

Cited by 79PDFScholar
2020

CP-GAN: Context Pyramid Generative Adversarial Network for Speech Enhancement

ICASSP 2020accepted

The topic of speech enhancement has been largely improved recently, especially with the development of generative adversarial networks (GANs). However prior methods simply follow the GAN architectures from computer vision tasks without specific designs for the speech enhancement according to the aud…

Cited by 0SourceScholar
2019

Graphonomy: Universal Human Parsing via Graph Transfer Learning

CVPR 2019poster

Prior highly-tuned human parsing models tend to fit towards each dataset in a specific domain or with discrepant label granularity, and can hardly be adapted to other human parsing tasks without extensive re-training. In this paper, we aim to learn a single universal human parsing model that can tac…

Cited by 228PDFcodeScholar
2019

Layout-Graph Reasoning for Fashion Landmark Detection

CVPR 2019poster

Detecting dense landmarks for diverse clothes, as a fundamental technique for clothes analysis, has attracted increasing research attention due to its huge application potential. However, due to the lack of modeling underlying semantic layout constraints among landmarks, prior works often detect amb…

Cited by 52PDFScholar
2018

Instance-level Human Parsing via Part Grouping Network

ECCV 2018poster

Instance-level human parsing towards real-world human analysis scenarios is still under-explored due to the absence of sufficient data resources and technical difficulty in parsing multiple instances in a single pass. Several related works all follow the ``parsing-by-detection" pipeline that heavily…

2018

Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis

NeurIPS 2018poster

Despite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most representative examples where the generation quality largely re…

Cited by 205SourcePDFScholar
2017

Look Into Person: Self-Supervised Structure-Sensitive Learning and a New Benchmark for Human Parsing

CVPR 2017poster

Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances and the coverage of challenging cases in unconstrained environment. In this pape…

Cited by 621PDFcodeScholar