← Search

Xiao Pu

9 accepted papers

2026

TGDD: Trajectory Guided Dataset Distillation with Balanced Distribution

AAAI 2026technical

Dataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature represent

Cited by 0SourcePDFScholar
2024

Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling

AAAI 2024technical

Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the whole dataset in practice. However, different selection subsets will lead to diff…

2024

Is Summary Useful or Not? An Extrinsic Human Evaluation of Text Summaries on Downstream Tasks

COLING 2024main

Research on automated text summarization typically uses human and automatic evaluation methods. While most recent studies focus on intrinsic evaluation, which assesses the general quality of summaries, e.g. coherence and informativeness, we concentrate on task-based extrinsic evaluation to determine…

2024

Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks

ACL 2024long

The widespread use of large language models (LLMs) is increasing the demand for methods that detect machine-generated text to prevent misuse. The goal of our study is to stress test the detectors’ robustness to malicious attacks under realistic scenarios. We comprehensively study the robustness of p…

2024

Style-Compress: An LLM-Based Prompt Compression Framework Considering Task-Specific Styles

EMNLP 2024finding

Prompt compression condenses contexts while maintaining their informativeness for different usage scenarios. It not only shortens the inference time and reduces computational costs during the usage of large language models, but also lowers expenses when using closed-source models. In a preliminary s…

2023

DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification

ICASSP 2023accepted

Very high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because…

Cited by 0SourceScholar
2023

FCIR: Rethink Aerial Image Super Resolution with Fourier Analysis

ICASSP 2023accepted

Recent years, deep-learning-based methods achieve remarkable improvements on the super-resolution (SR) task. However, recovering high-quality (HQ) texture from the low-quality (LQ) aerial image is still challenging due to the limited contextual modeling ability of current deep-learning methods as we…

Cited by 0SourceScholar
2023

On the Zero-Shot Generalization of Machine-Generated Text Detectors

EMNLP 2023short findings

The rampant proliferation of large language models, fluent enough to generate text indistinguishable from human-written language, gives unprecedented importance to the detection of machine-generated text. This work is motivated by an important research question: How will the detectors of machine-gen…

Cited by 0SourceScholar