← Search

Qing-Long Zhang

4 accepted papers

2024

Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight

NeurIPS 2024poster

This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales. This architecture not only leverages globa…

Cited by 3SourcePDFScholar
2024

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

ICML 2024poster

Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these c…

Cited by 18SourcePDFScholar
2023

M2TSR: Multi-Range and Mix-Grained Transformer for Single Image Super-Resolution

ICASSP 2023accepted

Recently, Transformers have shown impressive performance in image super-resolution (SR), due to exploiting strong representation ability of multi-head self-attention (MSA). However, existing methods typically calculate MSA in a single range and granularity, preventing the model from capturing suffic…

Cited by 0SourceScholar