← Search

Minfeng Zhu

10 accepted papers

2026

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

CVPR 2026

Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited

Cited by 0SourcecodeScholar
2026

Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Models

CVPR 2026

Chain-of-Thought (CoT) has recently shown encouraging progress in the vision language model. However, the pure-vision CoT (i.e., chain-of-vision) has been underexplored in visual in-context learning. In this paper, we introduce Diffusion Guided Chain-of-Vision, which integrates an explicit chain-of-

Cited by 0SourcecodeScholar
2026

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework

AAAI 2026technical

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep resear

Cited by 0SourcePDFScholar
2025

Don’t Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation

ACL 2025long

In this work, we investigate an important task named instruction-following text embedding, which generates dynamic text embeddings that adapt to user instructions, highlighting specific attributes of text. Despite recent advancements, existing approaches suffer from significant computational overhea…

2025

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

ICCV 2025poster

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reaso…

2024

Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

ACL 2024long

The surge in Large Language Models (LLMs) has revolutionized natural language processing, but fine-tuning them for specific tasks often encounters challenges in balancing performance and preserving general instruction-following abilities. In this paper, we posit that the distribution gap between tas…

2021

KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation

ICML 2021spotlight

Conventional unsupervised multi-source domain adaptation (UMDA) methods assume all source domains can be accessed directly. However, this assumption neglects the privacy-preserving policy, where all the data and computations must be kept decentralized. There exist three challenges in this scenario:…

2021

SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations

AAAI 2021technical

Semi-supervised variational autoencoders (VAEs) have obtained strong results, but have also encountered the challenge that good ELBO values do not always imply accurate inference results.In this paper, we investigate and propose two causes of this problem: (1) The ELBO objective cannot utilize the l…

2019

DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-To-Image Synthesis

CVPR 2019poster

In this paper, we focus on generating realistic images from text descriptions. Current methods first generate an initial image with rough shape and color, and then refine the initial image to a high-resolution one. Most existing text-to-image synthesis methods have two main problems. (1) These metho…

Cited by 797PDFScholar