← Search

Minh-Quan Le

4 accepted papers

2026

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

ICML 2026poster

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent…

Cited by 0SourceScholar
2025

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

ICLR 2025poster

While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Question Answering (VQA) and Human-Object Interaction (HOI) Reasoning, where it is critical to preserve scene attributes in…

2024

Learned Representation-Guided Diffusion Models for Large-Image Generation

CVPR 2024poster

To synthesize high-fidelity samples diffusion models typically require auxiliary data to guide the generation process. However it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed b…

2024

MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation

AAAI 2024technical

Few-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, kn…