← Search

Xiao Pan

4 accepted papers

2025

InsightEdit: Towards Better Instruction Following for Image Editing

CVPR 2025poster

In this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, InstructDiffusion, and SmartEdit have explored end-to-end editing. However, two limitations still remain: First, existing datasets suffer from low resolution, poor background consistency, and…

Cited by 2SourcePDFScholar
2023

Masked Audio Text Encoders are Effective Multi-Modal Rescorers

ACL 2023findings

Masked Language Models (MLMs) have proven to be effective for second-pass rescoring in Automatic Speech Recognition (ASR) systems. In this work, we propose Masked Audio Text Encoder (MATE), a multi-modal masked language model rescorer which incorporates acoustic representations into the input space…

2023

TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering

ICCV 2023poster

In this paper, we focus on the task of generalizable neural human rendering which trains conditional Neural Radiance Fields (NeRF) from multi-view videos of different characters. To handle the dynamic human motion, previous methods have primarily used a SparseConvNet (SPC)-based human representation…

Cited by 25PDFcodeScholar
2021

Contrastive Learning for Many-to-many Multilingual Neural Machine Translation

ACL 2021long

Existing multilingual machine translation approaches mainly focus on English-centric directions, while the non-English directions still lag behind. In this work, we aim to build a many-to-many translation system with an emphasis on the quality of non-English language directions. Our intuition is bas…