← Search

Zhuoran Yu

9 accepted papers

2026

Composition-Grounded Instruction Synthesis for Visual Reasoning

ICLR 2026poster

Pretrained multi-modal large language models (MLLMs) demonstrate strong performance on diverse multimodal tasks, but remain limited in reasoning capabilities for domains where annotations are difficult to collect. In this work, we focus on artificial image domains such as charts, rendered documents,…

Cited by 0SourcecodeScholar
2026

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

ICLR 2026poster

While Vision–language models (VLMs) have demonstrated remarkable performance across multi-modal tasks, their choice of vision encoders presents a fundamental weakness: their low-level features lack the robust structural and spatial information essential for document understanding and web agents. To…

Cited by 0SourceScholar
2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, overlooking realistic settings where numerical evidence in cha…

Cited by 0SourceScholar
2026

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge

ICML 2026poster

Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models in other modalities remains underexplored. In this work, we propose \LaViD—Language-to-Visual Knowledge Distillation—a simple and effective framework…

Cited by 0SourceScholar
2025

An Abnormal Audio Generation Method for Fault Diagnosis of Power Transformers

ICASSP 2025accepted

Existing deep learning-based models can achieve a prompt diagnosis of operational anomalies by analyzing the audios emitted from power transformers. However, the practical abnormal data are insufficient for model training, resulting in limited diagnostic performance. To address this problem, we prop…

Cited by 0SourceScholar
2025

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

ICCV 2025poster

Popular text-to-image (T2I) systems are trained on web-scraped data, which is heavily Amero and Euro-centric, underrepresenting the cultures of the Global South. To analyze these biases, we introduce CuRe, a novel and scalable benchmarking and scoring suite for cultural representativeness that lever…

2022

Group R-CNN for Weakly Semi-Supervised Object Detection With Points

CVPR 2022poster

We study the problem of weakly semi-supervised object detection with points (WSSOD-P), where the training data is combined by a small set of fully annotated images with bounding boxes and a large set of weakly-labeled images with only a single point annotated for each instance. The core of this task…

Cited by 55PDFcodeScholar
2020

Scale-Equalizing Pyramid Convolution for Object Detection

CVPR 2020poster

Feature pyramid has been an efficient method to extract features at different scales. Development over this method mainly focuses on aggregating contextual information at different levels while seldom touching the inter-level correlation in the feature pyramid. Early computer vision methods extracte…

Cited by 148PDFcodeScholar