← Search

Zhihong Zhu

57 accepted papers

2026

CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information Decoding

AAAI 2026technical

Medical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Model

Cited by 0SourcePDFScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizab

Cited by 0SourcePDFScholar
2025

$\text{D}_{2}\text{O}$: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

ICLR 2025poster

Efficient generative inference in Large Language Models (LLMs) is impeded by the growing memory demands of Key-Value (KV) cache, especially for longer sequences. Traditional KV Cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality…

Cited by 0SourcePDFScholar
2025

A Survey on Foundation Language Models for Single-cell Biology

ACL 2025long

The recent advancements in language models have significantly catalyzed progress in computational biology. A growing body of research strives to construct unified foundation models for single-cell biology, with language models serving as the cornerstone. In this paper, we systematically review the d…

Cited by 0SourcePDFScholar
2025

A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers

EMNLP 2025

Multi-modal intent recognition (MIR) requires integrating non-verbal cues from real-world contexts to enhance human intention understanding, which has attracted substantial research attention in recent years. Despite promising advancements, a comprehensive survey summarizing recent advances and new

2025

Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models

ACL 2025finding

Hallucination has emerged as a critical challenge for large language models (LLMs) and large vision-language models (LVLMs), particularly in high-stakes medical applications. Despite its significance, dedicated research on medical hallucination remains unexplored. In this survey, we first provide a…

Cited by 0SourcePDFScholar
2025

CellVerse: Do Large Language Models Really Understand Cell Biology?

NeurIPS 2025poster

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis ta…

Cited by 0SourcecodeScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

HTML: Hierarchical Topology Multi-task Learning for Semantic Parsing in Knowledge Base Question Answering

ACL 2025finding

Knowledge base question answering (KBQA) aims to answer natural language questions by reasoning over structured knowledge bases. Existing approaches often struggle with the complexity of mapping questions to precise logical forms, particularly when dealing with diverse entities and relations. In thi…

2025

Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation

AAAI 2025technical

Large Language Models (LLMs) demonstrate remarkable capabilities, yet struggle with hallucination and outdated knowledge when tasked with complex knowledge reasoning, resulting in factually incorrect outputs. Previous studies have attempted to mitigate it by retrieving factual knowledge from large-s…

2025

MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

NeurIPS 2025poster

While text-to-image models like GPT-4o-Image and FLUX are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it is crucial to align these models with desired behaviors based on fee…

Cited by 0SourcecodeScholar
2025

RTE-GMoE: A Model-agnostic Approach for Relation Triplet Extraction via Graph-based Mixture-of-Expert Mutual Learning

EMNLP 2025

Relation Triplet Extraction (RTE) is a fundamental while challenge task in knowledge acquisition, which identifies and extracts all triplets from unstructured text. Despite the recent advancements, the deep integration of the entity-, relation- and triplet-specific information remains a challenge. I

Cited by 0SourcePDFScholar
2025

Rethinking Decoding in Multi-intent Spoken Language Understanding

ICASSP 2025accepted

Multi-intent spoken language understanding (SLU) can handle multiple intent utterances in real-world scenarios, which has gained increasing research attention. Despite promising results achieved by existing joint models, they (1) perform utterance-level or token-level intent detection, resulting in…

Cited by 0SourceScholar
2025

Towards Zero-shot Cross-lingual SLU with Syntax-aware Multi-view Contrastive Learning

ICASSP 2025accepted

Recent state-of-the-art zero-shot cross-lingual spoken language understanding (SLU) models utilize contrastive learning to achieve multilingual semantics alignment between the original utterance and code-switched counterpart. Despite achieving promising results, we discover that they still suffer fr…

Cited by 0SourceScholar
2025

UniCoTT: A Unified Framework for Structural Chain-of-Thought Distillation

ICLR 2025poster

Chains of thought (CoTs) have achieved success in enhancing the reasoning capabilities of large language models (LLMs), while their effectiveness is predominantly observed in LLMs. Existing solutions methods adopt distillation to inject chain-of-thought capabilities into small models (SLMs). Howeve…

2025

VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

CVPR 2025poster

Large Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which significantly impedes their real-world usage. To alleviate VH, various decoding strategies have been proposed to enhance visual information. However, many of these…

2024

Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment

AAAI 2024technical

Multi-intent spoken language understanding (SLU) has garnered growing attention due to its ability to handle multiple intent utterances, which closely mirrors practical scenarios. Unlike traditional SLU, each intent in multi-intent SLU corresponds to its designated scope for slots, which occurs in…

2024

Alignment before Awareness: Towards Visual Question Localized-Answering in Robotic Surgery via Optimal Transport and Answer Semantics

COLING 2024main

The visual question localized-answering (VQLA) system has garnered increasing attention due to its potential as a knowledgeable assistant in surgical education. Apart from providing text-based answers, VQLA can also pinpoint the specific region of interest for better surgical scene understanding. Al…

2024

AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition

NAACL 2024long

Recent advancements in large language models (LLMs) have shown promise in multi-step reasoning tasks, yet their reliance on extensive manual labeling to provide procedural feedback remains a significant impediment. To address this challenge, in this paper, we propose a novel self-supervised framewor…

Cited by 25SourcePDFScholar
2024

Code-Switching Can be Better Aligners: Advancing Cross-Lingual SLU through Representation-Level and Prediction-Level Alignment

ACL 2024short

Zero-shot cross-lingual spoken language understanding (SLU) can promote the globalization application of dialog systems, which has attracted increasing attention. While current code-switching based cross-lingual SLU frameworks have shown promising results, they (i) predominantly utilize contrastive…

2024

Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding

ACL 2024findings

Owing to the scarcity of labeled training data, Spoken Language Understanding (SLU) is still a challenging task in low-resource languages. Therefore, zero-shot cross-lingual SLU attracts more and more attention. Contrastive learning is widely applied to explicitly align representations of similar se…

2024

Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent Detection

EMNLP 2024main

Multimodal intent detection is designed to leverage diverse modalities for a comprehensive understanding of user intentions in real-world scenarios, thus playing a critical role in modern task-oriented dialogue systems. Existing methods have made great progress in modal alignment and fusion, however…

Cited by 1SourcePDFScholar
2024

Exploiting Auxiliary Caption for Video Grounding

AAAI 2024technical

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, w…

Cited by 17SourcePDFScholar
2024

Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game Theory

EMNLP 2024main

Large Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which hinders their application in multimodal understanding and decision-making. In this work, we introduce a novel plug-and-play train-free decoding algorithm named Gam…

2024

InfoEnh: Towards Multimodal Sentiment Analysis via Information Bottleneck Filter and Optimal Transport Alignment

COLING 2024main

In recent years, Multimodal Sentiment Analysis (MSA) leveraging deep learning has demonstrated exceptional performance in a wide range of domains. Its success lies in effectively utilizing information from multiple modalities to analyze sentiments. Despite these advancements, MSA is confronted with…

Cited by 5SourcePDFScholar
2024

KC-Prompt: End-To-End Knowledge-Complementary Prompting for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continuous learning requires adapting quickly to incoming tasks while avoiding catastrophic forgetting. Typical solutions resort to a rehearsal buffer to replay old data, which is intractable to apply in real-world scenarios with limited memory and inaccessible privacy. Recently, with the emergence…

Cited by 0SourceScholar
2024

KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval

ECCV 2024poster

"Existing video-text retrieval methods predominantly focus on designing diverse cross-modal interaction mechanisms between captions and videos. However, those approaches diverge from human learning paradigms, where humans possess the capability to seek and associate knowledge from an open set, rathe…

Cited by 8SourcePDFScholar
2024

Knowledge-enhanced Prompt Tuning for Dialogue-based Relation Extraction with Trigger and Label Semantic

COLING 2024main

Dialogue-based relation extraction (DRE) aims to determine the semantic relation of a given pair of arguments from a piece of dialogue, which has received increasing attention. Due to the low information density of dialogue text, it is difficult for the model to focus on key information. To this end…

2024

LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

EMNLP 2024finding

Long-context Multimodal Large Language Models (MLLMs) demand substantial computational resources for inference as the growth of their multimodal Key-Value (KV) cache, in response to increasing input lengths, challenges memory and time efficiency. Unlike single-modality LLMs that manage only textual…

2024

Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System

EMNLP 2024finding

Due to the rapid development with pre-trained language models, fully end-to-end Task-Oriented Dialogue (TOD) systems exhibit superior performance. How to achieve the ability to efficiently retrieve entities in cross-domain large-scale databases is a key issue. Most existing end-to-end Task-Oriented…

Cited by 0SourcePDFScholar
2024

MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical Journey

NeurIPS 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding and generation, leading to their widespread adoption across various fields. Among these, the medical field is particularly well-suited for LLM applications, as many medical tasks can be enhanced by LLMs.…

Cited by 1SourcePDFScholar
2024

Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding

EMNLP 2024finding

The impressive capabilities of large language models (LLMs) have attracted extensive interests of applying LLMs to medical field. However, the complex nature of clinical environments presents significant hallucination challenges for LLMs, hindering their widespread adoption. In this paper, we addres…

2024

MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts

ACL 2024findings

As a crucial task in the task-oriented dialogue systems, spoken language understanding (SLU) has garnered increasing attention. However, errors from automatic speech recognition (ASR) often hinder the performance of understanding. To tackle this problem, we propose MoE-SLU, an ASR-Robust SLU framewo…

Cited by 2SourcePDFScholar
2024

Multi-perspective Improvement of Knowledge Graph Completion with Large Language Models

COLING 2024main

Knowledge graph completion (KGC) is a widely used method to tackle incompleteness in knowledge graphs (KGs) by making predictions for missing links. Description-based KGC leverages pre-trained language models to learn entity and relation representations with their names or descriptions, which shows…

2024

Relevance Is a Guiding Light: Relevance-aware Adaptive Learning for End-to-end Task-oriented Dialogue System

EMNLP 2024main

Retrieving accurate domain knowledge and providing helpful information are crucial in developing an effective end-to-end task-oriented dialogue system (E2ETOD). The field has witnessed numerous methods following the retrieve-then-generate paradigm and training their systems on one specific domain. H…

Cited by 1SourcePDFScholar
2024

TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing

IJCAI 2024poster

Multi-modal sarcasm detection (MSD), which aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, has garnered widespread attention. Recent approaches focus on designing sophisticated architectures or mechanisms to extract sarcastic cues from entire…

Cited by 10SourcePDFScholar
2024

Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport

AAAI 2024technical

Multi-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each inten…

2024

Towards Multi-modal Sarcasm Detection via Disentangled Multi-grained Multi-modal Distilling

COLING 2024main

Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, which has received increasing attention due to the rapid growth of multi-modal posts on modern social media. However, mainstream models process the input of each mo…

2024

UniMEEC: Towards Unified Multimodal Emotion Recognition and Emotion Cause

EMNLP 2024finding

Multimodal emotion recognition in conversation (MERC) and multimodal emotion-cause pair extraction (MECPE) have recently garnered significant attention. Emotions are the expression of affect or feelings; responses to specific events, or situations – known as emotion causes. Both collectively explain…

2024

What are the Generator Preferences for End-to-end Task-Oriented Dialog System?

EMNLP 2024main

Fully end-to-end task-oriented dialogue (EToD) systems have shown excellent performance, which requires the ability to retrieve entities accurately for generation. Existing methods improve the accuracy of entity retrieval and construct data flows between retrieval results and response generator, ach…

Cited by 0SourcePDFScholar
2024

Zero-Shot Spoken Language Understanding via Large Language Models: A Preliminary Study

COLING 2024main

Zero-shot Spoken Language Understanding (SLU) aims to enable task-oriented dialogue systems to understand user needs without training data. Challenging but worthwhile, zero-shot SLU reduces the time and effort that data labeling takes. Recent advancements in large language models (LLMs), such as GPT…

Cited by 14SourcePDFScholar
2023

A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding

ICASSP 2023accepted

Multi-intent detection and slot filling joint models are gaining increasing traction since they are closer to complicated real-world scenarios. However, existing approaches (1) focus on identifying implicit correlations between utterances and one-hot encoded labels in both tasks while ignoring expli…

Cited by 0SourceScholar
2023

Accelerating Multiple Intent Detection and Slot Filling via Targeted Knowledge Distillation

EMNLP 2023long findings

Recent non-autoregressive Spoken Language Understanding (SLU) models attracts increasing attention owing to the high inference speed. However, most of them still (1) suffer from the multi-modality problem since the prior knowledge about the reference is relatively poor during inference; (2) fail to…

Cited by 0SourceScholar
2023

Enhancing Code-Switching for Cross-lingual SLU: A Unified View of Semantic and Grammatical Coherence

EMNLP 2023short main

Despite the success of spoken language understanding (SLU) in high-resource languages, achieving similar performance in low-resource settings, such as zero-shot scenarios, remains challenging due to limited labeled training data. To improve zero-shot cross-lingual SLU, recent studies have explored c…

Cited by 0SourceScholar
2023

G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory

ICCV 2023oral

The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learning requires two key properties: (1) alignment of features of similar samples, and (2) uniformity of the induced distrib…

Cited by 52PDFScholar
2023

Improving Retrieval-Based Dialogue System Via Syntax-Informed Attention

ICASSP 2023accepted

Multi-turn response selection is a challenging task due to its high demands on efficient extraction of the matching features from abundant information provided by context utterances. Since incorporating syntactic information like dependency structures into neural models can promote a better understa…

Cited by 0SourceScholar
2023

MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language Understanding

EMNLP 2023long findings

Enhancing the robustness towards Automatic Speech Recognition (ASR) errors is of great importance for Spoken Language Understanding (SLU). Trending ASR-robust SLU systems have witnessed impressive improvements through global contrastive learning. However, although most ASR errors occur only at local…

Cited by 0SourceScholar
2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

ACL 2023findings

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impair the understanding performance and lead to error propagation. Although there are some attempts to address this problem…

2023

MRRL: Modifying the Reference via Reinforcement Learning for Non-Autoregressive Joint Multiple Intent Detection and Slot Filling

EMNLP 2023long findings

With the rise of non-autoregressive approach, some non-autoregressive models for joint multiple intent detection and slot filling have obtained the promising inference speed. However, most existing SLU models (1) suffer from the multi-modality problem that leads to reference intents and slots may no…

Cited by 0SourceScholar
2023

SSVMR: Saliency-Based Self-Training for Video-Music Retrieval

ICASSP 2023accepted

With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws much attention by research community. As other cross-modal learning tasks, existing VMR approaches usually attempt to m…

Cited by 0SourceScholar
2023

Syntax Matters: Towards Spoken Language Understanding via Syntax-Aware Attention

EMNLP 2023short findings

Spoken Language Understanding (SLU), a crucial component of task-oriented dialogue systems, has consistently garnered attention from both academic and industrial communities. Although incorporating syntactic information into models has the potential to enhance the comprehension of user utterances an…

Cited by 0SourceScholar
2023

Towards Unified Spoken Language Understanding Decoding via Label-aware Compact Linguistics Representations

ACL 2023findings

Joint intent detection and slot filling models have shown promising success in recent years due to the high correlations between the two tasks. However, previous works independently decode the two tasks, which could result in misaligned predictions for both tasks. To address this shortcoming, we pro…

2023

Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation

ICCV 2023poster

Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between im…

Cited by 42PDFScholar