← Search

Richard Yi Da Xu

9 accepted papers

2025

Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models

ACL 2025long

Interpretation is critical for disease diagnosis, but existing models struggle to balance predictive accuracy with human-understandable rationales. While large language models (LLMs) offer strong reasoning abilities, their clinical use is limited by high computational costs and restricted multimodal…

2025

ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series

ACL 2025finding

Large language models (LLMs) have shown remarkable performance in vision-language tasks, but their application in the medical field remains underexplored, particularly for integrating structured time series data with unstructured clinical notes. In clinical practice, dynamic time series data, such a…

Cited by 0SourcePDFScholar
2025

TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot Learning

IJCAI 2025

Compositional Zero-Shot Learning (CZSL) aims to recognize novel state-object compositions by leveraging the shared knowledge of their primitive components. Despite considerable progress, effectively calibrating the bias between semantically similar multimodal representations, as well as generalizing

2023

Domain Decorrelation with Potential Energy Ranking

AAAI 2023technical

Machine learning systems, especially the methods based on deep learning, enjoy great success in modern computer vision tasks under ideal experimental settings. Generally, these classic deep learning methods are built on the i.i.d. assumption, supposing the training and test data are drawn from the s…

2023

Robust Feature Rectification of Pretrained Vision Models for Object Recognition

AAAI 2023technical

Pretrained vision models for object recognition often suffer a dramatic performance drop with degradations unseen during training. In this work, we propose a RObust FEature Rectification module (ROFER) to improve the performance of pretrained models against degradations. Specifically, ROFER first es…

Cited by 0SourcePDFScholar
2021

Capturing Uncertainty in Unsupervised GPS Trajectory Segmentation Using Bayesian Deep Learning

AAAI 2021technical

Intelligent transportation management requires not only statistical information on users' mobility patterns, but also knowledge of their corresponding transportation modes. While GPS trajectories can be readily obtained from GPS sensors found in modern smartphones and vehicles, these massive geospat…

Cited by 25SourcePDFScholar
2021

On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization

IJCAI 2021poster

The prevailing thinking is that orthogonal weights are crucial to enforcing dynamical isometry and speeding up training. The increase in learning speed that results from orthogonal initialization in linear networks has been well-proven. However, while the same is believed to also hold for nonlinear…

Cited by 41SourcePDFScholar
2020

End-to-end Dynamic Matching Network for Multi-view Multi-person 3d Pose Estimation

ECCV 2020poster

As an important computer vision task, 3d human pose estimation in a multi-camera, multi-person setting has received widespread attention and many interesting applications have been derived from it. Traditional approaches use a 3d pictorial structure model to handle this task. However, these models s…

Cited by 50SourcePDFScholar