← Search

Peirou Liang

2 accepted papers

2026

ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs

AAAI 2026technical

Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we

Cited by 0SourcePDFScholar
2025

Efficient Data Labeling by Hierarchical Crowdsourcing with Large Language Models

COLING 2025main

Large language models (LLMs) have received lots of attention for their impressive performance in in-context dialogues and their potential to revolutionize service industries with a new business model, Model-as-a-Service (MaaS). Automated data labeling is a natural and promising service. However, lab…

Cited by 1SourcePDFScholar