← Search

Chengxin Zheng

4 accepted papers

2026

Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR

ICML 2026poster

Multimodal large language models (MLLMs) excel at high-level reasoning yet fail on OCR tasks where fine-grained visual details are compromised or misaligned. We identify an overlooked optimization issue in multi-layer feature fusion. Skip pathways introduce direct back-propagation paths from high-le…

Cited by 0SourceScholar
2026

MoEA-Net: Modality-Incremental Expert Aggregation Network for Retinal Prognostic Prediction

AAAI 2026technical

Automated analysis of temporal changes in multimodal retinal images is critical for the prognostic assessment of ophthalmic diseases. Unlike traditional single-timepoint diagnosis, tracking longitudinal changes across multiple imaging modalities introduces significant data bias challenges: (1) Imbal

Cited by 0SourcePDFScholar
2025

MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report Generation

AAAI 2025technical

The automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans involve extensive medical entities, such as diverse anatomy regions and lesions, exhibiting highly inconsistent spatial pat…

2024

See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning

EMNLP 2024finding

Brain CT report generation is significant to aid physicians in diagnosing cranial diseases.Recent studies concentrate on handling the consistency between visual and textual pathological features to improve the coherence of report.However, there exist some challenges: 1) Redundant visual representing…