← Search

Mengjie Wu

5 accepted papers

2026

PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection

ICML 2026poster

Large models (LMs) are powerful content generators, yet their open‑ended nature can also introduce potential risks, such as generating harmful or biased content. Existing guardrails mostly perform post-hoc detection that may expose unsafe content before it is caught, and the latency constraints furt…

Cited by 0SourceScholar
2025

Analogy-based Multi-Turn Jailbreak against Large Language Models

NeurIPS 2025poster

Large language models (LLMs) are inherently designed to support multi-turn interactions, which opens up new possibilities for jailbreak attacks that unfold gradually and potentially bypass safety mechanisms more effectively than single-turn attacks. However, current multi-turn jailbreak methods are…

Cited by 0SourceScholar
2025

Transfer Learning of Real Image Features with Soft Contrastive Loss for Fake Image Detection

AAAI 2025technical

In the last few years, the artifact patterns in fake images synthesized by different generative models have been inconsistent, leading to the failure of previous research that relied on spotting subtle differences between real and fake. In our preliminary experiments, we find that the artifacts in f…

Cited by 0SourcePDFScholar
2024

Learning-Based Efficient Phase- Amplitude Modulation and Hybrid Control for MRI-Guided Focused Ultrasound Treatment

RA-L 2024

Magnetic resonance-guided focused ultrasound (MRg-FUS) has become attractive, accredited to its non-invasive nature. However, ultrasound beams focusing and steering is still challenging owing to aberrations induced by soft tissue heterogeneity. In particular for beam motion control to ensure real-ti

Cited by 1SourceScholar
2024

TraceEvader: Making DeepFakes More Untraceable via Evading the Forgery Model Attribution

AAAI 2024technical

In recent few years, DeepFakes are posing serve threats and concerns to both individuals and celebrities, as realistic DeepFakes facilitate the spread of disinformation. Model attribution techniques aim at attributing the adopted forgery models of DeepFakes for provenance purposes and providing expl…

Cited by 7SourcePDFScholar