← Search

Minjie Hong

4 accepted papers

2026

DiagramGPT-Llama3:Enabling Editable, High-Fidelity Diagram Generation with Vision Large Language Models

AAAI 2026technical

The automation of diagram generation has gained significant attention in recent years. Previous studies mainly focused on generating diagrams from natural language, but often lacked support for user-friendly editing like drag-and-drop. This paper proposes a novel task: generating editable, high-fide

Cited by 0SourcePDFScholar
2026

Thinking with Programming Vision: Towards a Unified View for Thinking with Images

CVPR 2026

Multimodal large language models (MLLMs) that "think with images" can interactively use tools to reason about visual inputs, but current approaches often rely on a narrow set of tools with limited real-world necessity and scalability. In this work, we first reveal a critical and previously overlooke

Cited by 0SourcecodeScholar
2025

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025poster

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind…

Cited by 11SourcePDFScholar
2024

AudioVSR: Enhancing Video Speech Recognition with Audio Data

EMNLP 2024main

Visual Speech Recognition (VSR) aims to predict spoken content by analyzing lip movements in videos. Recently reported state-of-the-art results in VSR often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are insufficient compared to the audi…

Cited by 1SourcePDFScholar