← Search

Zhiyi Hou

3 accepted papers

2026

HEV Generative Sandbox: A Framework for Assessing Domain-Specific Social Risks Through Human-LLM Simulation

AAAI 2026technical

Deploying Large Language Models (LLMs) in specialized domains introduces significant societal and compliance risks, including bias amplification, misinformation propagation, and privacy violations. These risks predominantly emerge from the dynamic interactions between LLMs and humans in specific con

Cited by 0SourcePDFScholar
2024

Causality Based Front-door Defense Against Backdoor Attack on Language Models

ICML 2024poster

We have developed a new framework based on the theory of causal inference to protect language models against backdoor attacks. Backdoor attackers can poison language models with different types of triggers, such as words, sentences, grammar, and style, enabling them to selectively modify the decisio…

2024

Self-Powered LLM Modality Expansion for Large Speech-Text Models

EMNLP 2024main

Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities. Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable bene…