← Search

Huixuan Zhang

9 accepted papers

2025

DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language Models

EMNLP 2025

While large language models (LLMs) demonstrate remarkable capabilities across a wide range of tasks, they remain vulnerable to generating outputs that are potentially harmful. Red teaming, which involves crafting adversarial inputs to expose vulnerabilities, is a widely adopted approach for evaluati

2025

Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models

EMNLP 2025

In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating knowledge across these modalities during multimodal knowledge reasoning, leading

2025

ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs

ACL 2025long

Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination detection methods leveraging hidden states predominantly focus on static and isolated representations, overlooking their…

Cited by 0SourcePDFScholar
2025

MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency

ACL 2025finding

Multimodal large language models (MLLMs) are prone to non-factual or outdated knowledge issues, highlighting the importance of knowledge editing. Many benchmark has been proposed for researching multimodal knowledge editing. However, previous benchmarks focus on limited scenarios due to the lack of…

Cited by 0SourcePDFScholar
2025

R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models

EMNLP 2025

Text-to-image models frequently fail to achieve perfect alignment with textual prompts, particularly in maintaining proper semantic binding between semantic elements in the given prompt. Existing approaches typically require costly retraining or focus on only correctly generating the attributes of e

Cited by 0SourcePDFScholar
2025

Tracing Training Footprints: A Calibration Approach for Membership Inference Attacks Against Multimodal Large Language Models

EMNLP 2025

With the increasing scale of training data for Multimodal Large Language Models (MLLMs) and the lack of data details, there is growing concern about privacy breaches and data security issues. Under black-box access, exploring effective Membership Inference Attacks (MIA) has garnered increasing atten

Cited by 0SourcePDFScholar
2024

Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection

COLING 2024main

Hyperbole, or exaggeration, is a common linguistic phenomenon. The detection of hyperbole is an important part of understanding human expression. There have been several studies on hyperbole detection, but most of which focus on text modality only. However, with the development of social media, peop…

2024

PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models

EMNLP 2024finding

Large language models (LLMs) are known to be trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. This inclusion can lead to cheatingly high scores on model leaderboards, yet result in disappointing performance in real-world applicat…

Cited by 3SourcePDFScholar