← Search

Liqiang Niu

5 accepted papers

2026

UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings

ICLR 2026poster

The remarkable success of multimodal large language models (MLLMs) has driven advances in multimodal embeddings, yet existing models remain inherently discriminative, limiting their ability to benefit from reasoning-driven generation paradigm. In this work, we pioneer the exploration of generative e…

Cited by 0SourceScholar
2025

AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity

ACL 2025finding

Recently, large multimodal models (LMMs) have achieved significant advancements. When dealing with high-resolution images, dominant LMMs typically divide them into multiple local images and a global image, leading to a large number of visual tokens. In this work, we introduce AVG-LLaVA, an LMM that…

2025

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

EMNLP 2025

Universal multimodal embedding models play a critical role in tasks such as interleaved image-text retrieval, multimodal RAG, and multimodal clustering. However, our empirical results indicate that existing LMM-based embedding models trained with the standard InfoNCE loss exhibit a high degree of ov

Cited by 0SourcePDFScholar
2025

TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding

EMNLP 2025

Text-rich images are ubiquitous in real-world applications, serving as a critical medium for conveying complex information and facilitating accessibility.Despite recent advances driven by Multimodal Large Language Models (MLLMs), existing benchmarks suffer from limited scale, fragmented scenarios, a

Cited by 0SourcePDFScholar
2024

UMTIT: Unifying Recognition, Translation, and Generation for Multimodal Text Image Translation

COLING 2024main

Prior research in Image Machine Translation (IMT) has focused on either translating the source image solely into the target language text or exclusively into the target image. As a result, the former approach lacked the capacity to generate target images, while the latter was insufficient in produci…