← Search

Yuming Qiao

4 accepted papers

2026

Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning Learning

AAAI 2026technical

Spatial understanding is a critical capability for LVLMs (Large Vision-Language Models) to advance embodied AI applications. Existing works primarily focus on enhancing spatial understanding within a single frame, i.e., injecting 3D spatial concepts into LVLMs under single coordinate system. However

Cited by 0SourcePDFScholar
2026

JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning

CVPR 2026

With the advancement of Vision-Language Models (VLMs), employing VLM-as-a-Judge for visual evaluation has become a widely adopted metric in vision research. However, existing VLM-as-a-Judge approaches suffer from biased scoring outcomes with low discrimination and lack the capacity for unified multi

Cited by 0SourcecodeScholar
2025

MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAG

IJCAI 2025

Language-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) m

Cited by 0SourcePDFScholar
2024

BARET: Balanced Attention Based Real Image Editing Driven by Target-Text Inversion

AAAI 2024technical

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple conditions (e.g., mask, sketch, caption), and time consuming fine-t…

Cited by 5SourcePDFScholar