← Search

Huang Chen

4 accepted papers

2024

FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models

NeurIPS 2024poster

Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structured content. Rendering algorithms decompose and simulate the imaging process of a camera, while are limited by the accura…

Cited by 1SourcePDFScholar
2024

Visual Hallucination Elevates Speech Recognition

AAAI 2024technical

Due to the detrimental impact of noise on the conventional audio speech recognition (ASR) task, audio-visual speech recognition~(AVSR) has been proposed by incorporating both audio and visual video signals. Although existing methods have demonstrated that the aligned visual input of lip movements ca…

Cited by 5SourcePDFScholar
2023

Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration

ICCV 2023poster

We propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including document analysis, retrieval, and office automation. Unlike state-of-the-art approaches that rely on multi-stage techni…

Cited by 15PDFScholar
2022

CoCGAN: Contrastive Learning for Adversarial Category Text Generation

COLING 2022main

The task of generating texts of different categories has attracted more and more attention in the area of natural language generation recently. Meanwhile, generative adversarial net (GAN) has demonstrated its effectiveness on text generation, and is further applied to category text generation in lat…