← Search

Spencer Hong

2 accepted papers

2026

EPSVec: Efficient and Private Synthetic Text Generation via Dataset Vectors

ICML 2026poster

High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have emerged as powerful engines for generating it. However, exis…

Cited by 0SourceScholar
2025

Watermark under Fire: A Robustness Evaluation of LLM Watermarking

EMNLP 2025

Various watermarking methods (“watermarkers”) have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions remain under-explored: i) What are the strengths/limitations of various watermarkers, especially their attack robustness? ii