← Search

Zhu Zhang

10 accepted papers

2026

A MEDICAL MULTIMODAL DIAGNOSTIC FRAMEWORK INTEGRATING VISION-LANGUAGE MODELS AND LOGIC TREE REASONING

ICASSP 2026poster

With the rapid growth of large language models (LLMs) and vision-language models (VLMs) in medicine, simply integrating clinical text and medical imaging does not guarantee reliable reasoning. Existing multimodal models often produce hallucinations or inconsistent chains of thought, limiting clinica…

Cited by 0SourcePDFScholar
2026

MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models

AAAI 2026technical

Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-gr

Cited by 0SourcePDFScholar
2024

Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models

NAACL 2024findings

Fact-checking is an essential task in NLP that is commonly utilized to validate the factual accuracy of a piece of text. Previous approaches mainly involve the resource-intensive process of fine-tuning pre-trained language models on specific datasets. In addition, there is a notable gap in datasets…

2021

Learning to Rehearse in Long Sequence Memorization

ICML 2021spotlight

Existing reasoning tasks often have an important assumption that the input contents can be always accessed while reasoning, requiring unlimited storage resources and suffering from severe time delay on long sequences. To achieve efficient reasoning on long sequences with limited storage resources, m…

Cited by 12SourcePDFScholar
2021

RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems

ACL 2021long

For task-oriented dialog systems to be maximally useful, it must be able to process conversations in a way that is (1) generalizable with a small number of training examples for new task domains, and (2) robust to user input in various styles, modalities, or domains. In pursuit of these goals, we in…

Cited by 49SourcePDFScholar
2021

UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis

NeurIPS 2021poster

Conditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as well as their combinations. In this paper, instead of investigating these control signals separately, we propose a new t…

Cited by 77SourcePDFScholar
2020

Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language Grounding

NeurIPS 2020poster

Weakly-supervised vision-language grounding aims to localize a target moment in a video or a specific region in an image according to the given sentence query, where only video-level or image-level sentence annotations are provided during training. Most existing approaches employ the MIL-based or re…

Cited by 147SourcePDFScholar
2020

Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding

IJCAI 2020poster

Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper, we explore spatio-temporal video grounding on unaligned data…

Cited by 0SourcePDFScholar
2020

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

CVPR 2020poster

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depicting an object, STVG aims to localize the spatio-temporal tube of the queried object. STVG has two challenging settings: (1…

Cited by 134PDFcodeScholar