← Search

Jiachen Ma

4 accepted papers

2026

PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs Hallucination

AAAI 2026technical

Large Language Models (LLMs) perform excellently in fake news detection tasks, but their outputs are often accompanied by hallucinations, i.e., generated content that is contradictory to facts. Previous studies have mostly mitigated hallucinations through prompt design. However, this paper reveals t

Cited by 0SourcePDFScholar
2026

Reflector: Internalizing Step-wise Reflection against Indirect Jailbreaks

ICML 2026poster

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process. To address these vulnerabilities, we propose Refle…

Cited by 0SourceScholar
2025

Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

NAACL 2025findings

Text-to-image (T2I) models can be maliciously used to generate harmful content such as sexually explicit, unfaithful, and misleading or Not-Safe-for-Work (NSFW) images. Previous attacks largely depend on the availability of the diffusion model or involve a lengthy optimization process. In this work,…

2023

CAME: Contrastive Automated Model Evaluation

ICCV 2023poster

The Automated Model Evaluation (AutoEval) framework entertains the possibility of evaluating a trained machine learning model without resorting to a labeled testing set. Despite the promise and some decent results, the existing AutoEval methods heavily rely on computing distribution shifts between…

Cited by 9PDFcodeScholar