← Search

Munan Ning

9 accepted papers

2026

WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

ICML 2026poster

Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and…

Cited by 0SourceScholar
2025

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step

NeurIPS 2025poster

Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled from layout planning, making it difficult to refine the layo…

Cited by 0SourceScholar
2025

UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation mechanisms face limitations due to the significant human workload required to desi…

Cited by 0SourcePDFScholar
2024

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

ICLR 2024poster

The video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N ≥ 3) beyond vision and language. We thus propose LanguageBind, taking the language as the bind…

2024

Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable Repainting

ECCV 2024poster

"Recent image-to-3D methods achieve impressive results with plausible 3D geometry due to the development of diffusion models and optimization techniques. However, existing image-to-3D methods suffer from texture deficiencies in novel views, including multi-view inconsistency and quality degradation.…

Cited by 27SourcePDFScholar
2024

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

EMNLP 2024main

Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tok…

2022

Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation

ECCV 2022poster

"Research into Few-shot Semantic Segmentation (FSS) has attracted great attention, with the goal to segment target objects in a query image given only a few annotated support images of the target class. A key to this challenging task is to fully utilize the information in the support images by explo…

2021

Multi-Anchor Active Domain Adaptation for Semantic Segmentation

ICCV 2021poster

Unsupervised domain adaption has proven to be an effective approach for alleviating the intensive workload of manual annotation by aligning the synthetic source-domain data and the real-world target-domain samples. Unfortunately, mapping the target-domain distribution to the source-domain unconditio…

Cited by 60PDFcodeScholar
2020

Hierarchical Clustering With Hard-Batch Triplet Loss for Person Re-Identification

CVPR 2020poster

For clustering-guided fully unsupervised person reidentification (re-ID) methods, the quality of pseudo labels generated by clustering directly decides the model performance. In order to improve the quality of pseudo labels in existing methods, we propose the HCT method which combines hierarchical c…

Cited by 368PDFcodeScholar