← Search

Zhoujian Sun

3 accepted papers

2026

anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding

AAAI 2026technical

The advent of multimodal large language models (MLLMs) has sparked interest in their application to electrocardiogram (ECG) analysis. However, existing ECG-focused MLLMs primarily focus on report generation tasks, often limited to single 12-lead, short-duration (10s) ECG inputs, thereby underutilizi

Cited by 0SourcePDFScholar
2024

Multi-Modality Speech Recognition Driven by Background Visual Scenes

ICASSP 2024accepted

Visual information is often used as a complementary cue for automatic speech recognition in noisy environments. Most previous studies utilize visual information of target speakers (e.g., lip movements) to improve the recognition performance of audio-visual speech recognition (AVSR) models. However,…

Cited by 0SourceScholar
2022

On Tracking Dialogue State by Inheriting Slot Values in Mentioned Slot Pools

IJCAI 2022poster

Dialogue state tracking (DST) is a component of the task oriented dialogue system. It is responsible for extracting and managing slots, where each slot represents a part of the information to accomplish a task, and slot value is updated recurrently in each dialogue turn. However, many DST models can…