← Search

Li Zhou

30 accepted papers

2026

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

ICLR 2026poster

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks t…

Cited by 0SourcecodeScholar
2026

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

IJCAI 2026

Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic--acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as

Cited by 0Scholar
2026

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

ICASSP 2026poster

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external…

Cited by 0SourcePDFScholar
2025

A Compressive Memory-based Retrieval Approach for Event Argument Extraction

COLING 2025main

Recent works have demonstrated the effectiveness of retrieval augmentation in the Event Argument Extraction (EAE) task. However, existing retrieval-based EAE methods have two main limitations: (1) input length constraints and (2) the gap between the retriever and the inference model. These issues li…

Cited by 3SourcePDFScholar
2025

Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge

NAACL 2025long

Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet often lack a robust methodology to dissect these phenomena comprehensively. Our work aims to bridge this gap by delving into the Food domain—a universally relevant yet culturally diverse aspect of hu…

2025

Engage for All: Making Ordinary Image Descriptions Appealing Again!

ICCV 2025poster

In recent years, multi-modal large language models (MLLMs) have been successfully adopted to generate humorous and engaging descriptions for internet memes. While, it is challenging for the same approaches to apply to ordinary images which lack of inherent funny or exaggerated contents. Thus, crafti…

2025

Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction Modeling

ICASSP 2025accepted

Document-level relation extraction aims at extracting relational facts between two entities in a document. Existing approaches mainly focus on target entities, utilizing techniques such as graph neural networks to enhance their representations. However, they ignore the rich semantic correlations amo…

Cited by 0SourceScholar
2025

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

EMNLP 2025

The human-centered word association test (WAT) serves as a cognitive proxy, revealing sociocultural variations through culturally shared semantic expectations and implicit linguistic patterns shaped by lived experiences. We extend this test into an LLM-adaptive, free-relation task to assess the alig

2025

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

EMNLP 2025

Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions. To bridge this gap, we introduce Hanf

Cited by 0SourcePDFScholar
2025

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

ICCV 2025poster

Hallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepanc…

2025

Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model

AAAI 2025technical

Large Multimodal Models (LMMs) have significantly progressed by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation by integrating segmentation models. Despite the innovations, existing works’ t…

2025

Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles

ACL 2025long

User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and…

2025

RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions

EMNLP 2025

Retrieval-Augmented Generation (RAG) has emerged as a key paradigm for enhancing large language models by incorporating external knowledge. However, current RAG methods exhibit limited capabilities in complex RAG scenarios and suffer from limited task diversity. To address these limitations, we prop

2025

W-ControlUDA: Weather-Controllable Diffusion-assisted Unsupervised Domain Adaptation for Semantic Segmentation

RA-L 2025

Image generation has emerged as a potent strategy to enrich training data for unsupervised domain adaptation (UDA) of semantic segmentation in adverse weathers due to the scarcity of labelled target domain data. Previous UDA works commonly utilize generative adversarial networks (GANs) to translate

Cited by 6SourceScholar
2024

AdaFL: Adaptive Client Selection and Dynamic Contribution Evaluation for Efficient Federated Learning

ICASSP 2024accepted

Federated learning is a collaborative machine learning framework where multiple clients jointly train a global model. To mitigate communication overhead, it is common to select a subset of clients for participation in each training round. However, existing client selection strategies often rely on a…

Cited by 0SourceScholar
2024

Beyond Single-Event Extraction: Towards Efficient Document-Level Multi-Event Argument Extraction

ACL 2024findings

Recent mainstream event argument extraction methods process each event in isolation, resulting in inefficient inference and ignoring the correlations among multiple events. To address these limitations, here we propose a multiple-event argument extraction model DEEIA (Dependency-guided Encoding and…

2024

ESCoT: Towards Interpretable Emotional Support Dialogue Systems

ACL 2024long

Understanding the reason for emotional support response is crucial for establishing connections between users and emotional support dialogue systems. Previous works mostly focus on generating better responses but ignore interpretability, which is extremely important for constructing reliable dialogu…

2024

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

EMNLP 2024main

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQA, a manually curated, fine-grained image-text dataset capturing the intricate f…

2024

MLPs Compass: What is Learned When MLPs are Combined with PLMs?

ICASSP 2024accepted

While Transformer-based pre-trained language models and their variants exhibit strong semantic representation capabilities, the question of comprehending the information gain derived from the additional components of PLMs remains an open question in this field. Motivated by recent efforts that prove…

Cited by 0SourceScholar
2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2023

Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural Features

EMNLP 2023long findings

The increasing ubiquity of language technology necessitates a shift towards considering cultural diversity in the machine learning realm, particularly for subjective tasks that rely heavily on cultural nuances, such as Offensive Language Detection (OLD). Current understanding underscores that these…

Cited by 0SourcecodeScholar
2023

Joint Visual Grounding and Tracking With Natural Language Specification

CVPR 2023poster

Tracking by natural language specification aims to locate the referred target in a sequence based on the natural language description. Existing algorithms solve this issue in two steps, visual grounding and tracking, and accordingly deploy the separated grounding model and tracking model to implemen…

2023

Substructure Aware Graph Neural Networks

AAAI 2023technical

Despite the great achievements of Graph Neural Networks (GNNs) in graph learning, conventional GNNs struggle to break through the upper limit of the expressiveness of first-order Weisfeiler-Leman graph isomorphism test algorithm (1-WL) due to the consistency of the propagation paradigm of GNNs with…

2021

Generating Self-Contained and Summary-Centric Question Answer Pairs via Differentiable Reward Imitation Learning

EMNLP 2021main

Motivated by suggested question generation in conversational news recommendation systems, we propose a model for generating question-answer pairs (QA pairs) with self-contained, summary-centric questions and length-constrained, article-summarizing answers. We begin by collecting a new dataset of new…

2018

Learning for Disparity Estimation Through Feature Constancy

CVPR 2018poster

Stereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four steps, or use different networks to deal with different step…