← Search

Mengyue Wu

38 accepted papers

2026

A Human-Centric Pipeline for Aligning Large Language Models with Chinese Medical Ethics

AAAI 2026technical

Recent advances in large language models (LLMs) have enabled their application to a range of healthcare tasks. However, aligning LLMs with the nuanced demands of medical ethics, especially under complex real-world scenarios, remains underexplored. In this work, we present MedES, a dynamic, scenario-

Cited by 0SourcePDFScholar
2026

From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity

ICLR 2026poster

Psychiatric comorbidity is clinically significant yet challenging due to the complexity of multiple co-occurring disorders. To address this, we develop a novel approach integrating synthetic patient electronic medical record (EMR) construction and multi-agent diagnostic dialogue generation. We creat…

Cited by 0SourceScholar
2026

PICOAUDIO2: TEMPORAL CONTROLLABLE TEXT-TO-AUDIO GENERATION WITH NATURAL LANGUAGE DESCRIPTION

ICASSP 2026oral

While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and input format. These models often suffer from poor audio quality in real datasets due to sole reliance on synthetic data. M…

Cited by 0SourcePDFScholar
2025

A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

NAACL 2025long

Large Vision-Language Models (LVLMs), despite their recent success, are hardly comprehensively tested for their cognitive abilities. Inspired by the prevalent use of the Cookie Theft task in human cognitive tests, we propose a novel evaluation benchmark to evaluate high-level cognitive abilities of…

Cited by 3SourcePDFScholar
2025

A Diverse and Effective Retrieval-Based Debt Collection System with Expert Knowledge

NAACL 2025industry

Designing effective debt collection systems is crucial for improving operational efficiency and reducing costs in the financial industry. However, the challenges of maintaining script diversity, contextual relevance, and coherence make this task particularly difficult. This paper presents a debt col…

Cited by 0SourcePDFScholar
2025

MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics

NAACL 2025industry

Large language models (LLMs) demonstrate significant potential in advancing medical applications, yet their capabilities in addressing medical ethics challenges remain underexplored. This paper introduces MedEthicEval, a novel benchmark designed to systematically evaluate LLMs in the domain of medic…

2025

PicoAudio: Enabling Precise Temporal Controllability in Text-to-Audio Generation

ICASSP 2025accepted

Recently, audio generation tasks have attracted considerable research interests. Despite rapid advancements in generating high-fidelity audio that is coarsely aligned with the text description, precise temporal controllability is still a challenge, which is essential to integrate audio generation wi…

Cited by 0SourceScholar
2025

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

ICASSP 2025accepted

The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and sy…

Cited by 0SourceScholar
2025

Toward Automatic Discovery of a Canine Phonetic Alphabet

ACL 2025long

Dogs communicate intelligently but little is known about the phonetic properties of their vocalization communication. For the first time, this paper presents an iterative algorithm inspired by human phonetic discovery, which is based on minimal pairs that determine phonemes by distinguishing differe…

Cited by 0SourcePDFScholar
2025

Tracking Life’s Ups and Downs: Mining Life Events from Social Media Posts for Mental Health Analysis

ACL 2025long

Social media platforms possess considerable potential in the realm of exploring mental health. Previous research has indicated that major life events can greatly impact individuals’ mental health. However, due to the complexity and ambiguity nature of life events, shedding its light on social media…

Cited by 0SourcePDFScholar
2025

WritingBench: A Comprehensive Benchmark for Generative Writing

NeurIPS 2025poster

Recent advancements in large language models (LLMs) have significantly enhanced text generation capabilities, yet evaluating their performance in generative writing remains a challenge. Existing benchmarks primarily focus on generic text generation or limited in writing tasks, failing to capture the…

Cited by 0SourcecodeScholar
2024

A Detailed Audio-Text Data Simulation Pipeline Using Single-Event Sounds

ICASSP 2024accepted

Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such descriptions are significantly limited. In this paper, we f…

Cited by 0SourceScholar
2024

Automatic Reconstruction of Ancient Chinese Pronunciations

EMNLP 2024finding

Reconstructing ancient Chinese pronunciation is a challenging task due to the scarcity of phonetic records. Different from historical linguistics’ comparative approaches, we reformulate this problem into a temporal prediction task with masked language models, digitizing existing phonology rules into…

2024

Mapping Long-term Causalities in Psychiatric Symptomatology and Life Events from Social Media

NAACL 2024long

Social media is a valuable data source for exploring mental health issues. However, previous studies have predominantly focused on the semantic content of these posts, overlooking the importance of their temporal attributes, as well as the evolving nature of mental disorders and symptoms.In this pap…

Cited by 0SourcePDFScholar
2024

Phonetic and Lexical Discovery of Canine Vocalization

EMNLP 2024finding

This paper attempts to discover communication patterns automatically within dog vocalizations in a data-driven approach, which breaks the barrier previous approaches that rely on human prior knowledge on limited data. We present a self-supervised approach with HuBERT, enabling the accurate classific…

Cited by 8SourcePDFScholar
2023

Detection of Multiple Mental Disorders from Social Media with Two-Stream Psychiatric Experts

EMNLP 2023long main

Existing Mental Disease Detection (MDD) research largely studies the detection of a single disorder, overlooking the fact that mental diseases might occur in tandem. Many approaches are not backed by domain knowledge (e.g., psychiatric symptoms) and thus fail to produce interpretable results. To ta…

Cited by 0SourceScholar
2023

Semantic Space Grounded Weighted Decoding for Multi-Attribute Controllable Dialogue Generation

EMNLP 2023long main

Controlling chatbot utterance generation with multiple attributes such as personalities, emotions and dialogue acts is a practically useful but under-studied problem. We propose a novel framework called DASC that possesses strong controllability with a weighted decoding paradigm, while improving…

Cited by 0SourcecodeScholar
2023

Transcribing Vocal Communications of Domestic Shiba lnu Dogs

ACL 2023findings

How animals communicate and whether they have languages is a persistent curiosity of human beings. However, the study of animal communications has been largely restricted to data from field recordings or in a controlled environment, which is expensive and limited in scale and variety. In this paper,…

2022

Can Audio Captions Be Evaluated With Image Caption Metrics?

ICASSP 2022accepted

Automated audio captioning aims at generating textual descriptions for an audio clip. To evaluate the quality of generated audio captions, previous works directly adopt image captioning metrics like SPICE and CIDEr, without justifying their suitability in this new domain, which may mislead the devel…

Cited by 0SourceScholar
2022

Category-Adapted Sound Event Enhancement with Weakly Labeled Data

ICASSP 2022accepted

Previous audio enhancement training usually requires clean signals with additive noises; hence commonly focuses on speech enhancement, where clean speech is easy to access. This paper goes beyond a broader sound event enhancement by using a weakly supervised approach via sound event detection (SED)…

Cited by 0SourceScholar
2022

D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented Chat

EMNLP 2022main

In a depression-diagnosis-directed clinical session, doctors initiate a conversation with ample emotional support that guides the patients to expose their symptoms based on clinical diagnosis criteria. Such a dialogue system is distinguished from existing single-purpose human-machine dialog systems,…

2022

Psychiatric Scale Guided Risky Post Screening for Early Detection of Depression

IJCAI 2022poster

Depression is a prominent health challenge to the world, and early risk detection (ERD) of depression from online posts can be a promising technique for combating the threat. Early depression detection faces the challenge of efficiently tackling streaming data, balancing the tradeoff between timelin…

2022

Symptom Identification for Interpretable Detection of Multiple Mental Disorders on Social Media

EMNLP 2022main

Mental disease detection (MDD) from social media has suffered from poor generalizability and interpretability, due to lack of symptom modeling. This paper introduces PsySym, the first annotated symptom identification corpus of multiple psychiatric disorders, to facilitate further research progress.…

2021

Building Interpretable Interaction Trees for Deep NLP Models

AAAI 2021technical

This paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituent…

Cited by 43SourcePDFScholar
2021

Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning

ICASSP 2021accepted

Automated audio captioning (AAC) aims at generating summarizing descriptions for audio clips. Multitudinous concepts are described in an audio caption, ranging from local information such as sound events to global information like acoustic scenery. Currently, the mainstream paradigm for AAC is the e…

Cited by 0SourceScholar
2021

Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events

ICASSP 2021accepted

Automated Audio Captioning is a cross-modal task, generating natural language descriptions to summarize the audio clips’ sound events. However, grounding the actual sound events in the given audio based on its corresponding caption has not been investigated. This paper contributes an Audio-Grounding…

Cited by 0SourceScholar
2020

Multiple Sound Sources Localization from Coarse to Fine

ECCV 2020poster

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework that disentangles audio and visual representations of different…