← Search

Jiang Gui

8 accepted papers

2026

HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics Modeling

ICML 2026poster

Spatial transcriptomics (ST) links gene expression with tissue morphology but remains expensive and low-throughput, motivating surrogates that infer expression from routine histology. Whole-slide H&E-to-ST inference pairs a gigapixel image with gene measurements at a sparse, irregular set of locatio…

Cited by 0SourceScholar
2026

Variation in Verification: Understanding Verification Dynamics in Large Language Models

ICLR 2026poster

Recent advances have shown that scaling test-time computation enables large language models (LLMs) to solve increasingly complex problems across diverse domains. One effective paradigm for test-time scaling (TTS) involves LLM generators producing multiple solution candidates, with LLM verifiers asse…

Cited by 0SourceScholar
2025

Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models

EMNLP 2025

Large Language Models (LLMs) offer transformative potential across diverse fields, yet their safe and effective deployment is hindered by inherent knowledge conflicts—stemming from temporal evolution, divergent sources, and contradictory guidelines. This challenge is particularly acute in medicine,

Cited by 0SourcePDFScholar
2025

Learning Sparsity for Effective and Efficient Music Performance Question Answering

ACL 2025short

Music performances, characterized by dense and continuous audio as well as seamless audio-visual integration, present unique challenges for multimodal scene understanding and reasoning. Recent Music Performance Audio-Visual Question Answering (Music AVQA) datasets have been proposed to reflect these…

Cited by 0SourcePDFScholar
2025

ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering

EMNLP 2025

Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not only accurate answers but also explanations that humans can easily understand an

Cited by 0SourcePDFScholar
2025

SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models

EMNLP 2025

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic approach that involves a capable base model, high-quality rea

2025

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

NAACL 2025findings

Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended tempora…

2024

Learning Musical Representations for Music Performance Question Answering

EMNLP 2024finding

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the audio-video QA demonstrate impressive capabilities on genera…