← Search

Zhouqiang Jiang

5 accepted papers

2026

Multimodal Fusion via Self-Consistent Task-Gradient Fields

ICML 2026poster

Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature extractors. Aggressively merging modalities entangles their representations, making the feature extractors fragile to in…

Cited by 0SourceScholar
2025

Putting People in LLMs’ Shoes: Generating Better Answers via Question Rewriter

AAAI 2025technical

Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermined by the vagueness of user questions. To address this issue, we introduce single-round instance-level prompt optimizat…

2025

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

COLING 2025main

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are unable to automatically identify such grouping, we argue that current VrDU approac…

2025

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

ICCV 2025poster

The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language models (MLLMs) exhibit impressing multimodal capabilities, they often fail in rarely encountered domain-specific tasks due…

2024

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex ta…