← Search

Junchao Liao

4 accepted papers

2026

LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

ICLR 2026poster

Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them a…

Cited by 0SourcecodeScholar
2025

Tora: Trajectory-oriented Diffusion Transformer for Video Generation

CVPR 2025poster

Recent advancements in Diffusion Transformer (DiT) have demonstrated remarkable proficiency in producing high-quality video content. Nonetheless, the potential of transformer-based diffusion models for effectively generating videos with controllable motion remains an area of limited exploration. Thi…

2025

TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video Synthesis

ICASSP 2025accepted

Recent developments in Video Diffusion Models (VDMs) have demonstrated remarkable capability to generate high-quality video content. Nonetheless, the potential of VDMs for creating transparent videos remains largely uncharted. In this paper, we introduce TransVDM, the first diffusion-based model spe…

Cited by 0SourceScholar
2022

Knowledge Mining With Scene Text for Fine-Grained Recognition

CVPR 2022poster

Recently, the semantics of scene text has been proven to be essential in fine-grained image classification. However, the existing methods mainly exploit the literal meaning of scene text for fine-grained recognition, which might be irrelevant when it is not significantly related to objects/scenes. W…

Cited by 20PDFcodeScholar