← Search

Siqi Ma

8 accepted papers

2026

A MEDICAL MULTIMODAL DIAGNOSTIC FRAMEWORK INTEGRATING VISION-LANGUAGE MODELS AND LOGIC TREE REASONING

ICASSP 2026poster

With the rapid growth of large language models (LLMs) and vision-language models (VLMs) in medicine, simply integrating clinical text and medical imaging does not guarantee reliable reasoning. Existing multimodal models often produce hallucinations or inconsistent chains of thought, limiting clinica…

Cited by 0SourcePDFScholar
2026

MPAS: Breaking Sequential Constraints of Multi-Agent Communication Topologies via Individual-Epistemic Message Propagation

AAAI 2026technical

Large language model (LLM)-driven agents are designed to handle a wide range of tasks autonomously. As tasks become increasingly composite, the integration of multiple agents into a graph-structured system offers a promising solution. Recent advances mainly architect the communication order among ag

Cited by 0SourcePDFScholar
2026

MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models

AAAI 2026technical

Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-gr

Cited by 0SourcePDFScholar
2025

Visual Agents as Fast and Slow Thinkers

ICLR 2025poster

Achieving human-level intelligence requires refining cognitive distinctions between \textit{System 1} and \textit{System 2} thinking. While contemporary AI, driven by large language models, demonstrates human-like traits, it falls short of genuine cognition. Transitioning from structured benchmarks…

2024

Random Entangled Tokens for Adversarially Robust Vision Transformer

CVPR 2024poster

Vision Transformers (ViTs) have emerged as a compelling alternative to Convolutional Neural Networks (CNNs) in the realm of computer vision showcasing tremendous potential. However recent research has unveiled a susceptibility of ViTs to adversarial attacks akin to their CNN counterparts. Adversaria…

Cited by 3SourcePDFScholar
2023

Private Image Generation With Dual-Purpose Auxiliary Classifier

CVPR 2023highlight

Privacy-preserving image generation has been important for segments such as medical domains that have sensitive and limited data. The benefits of guaranteed privacy come at the costs of generated images' quality and utility due to the privacy budget constraints. The utility is currently measured by…

Cited by 4SourcePDFScholar
2023

TransFlow: Transformer As Flow Learner

CVPR 2023highlight

Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based method…

Cited by 99SourcePDFScholar