← Search

Hailin Hu

9 accepted papers

2026

GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection

AAAI 2026technical

The rapid advancement of video generation models has made it increasingly challenging to distinguish AI-generated videos from real ones. This issue underscores the urgent need for effective AI-generated video detectors to prevent the dissemination of false information via such videos. However, the d

Cited by 0SourcePDFScholar
2026

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

ICLR 2026poster

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks, yet often suffer from inefficiencies due to redundant visual tokens. Existing token merging methods reduce sequence length but frequently disrupt spatial layouts and temporal continuity by disregardin…

Cited by 0SourcecodeScholar
2025

Data-Efficient Selection via Grammatical Complexity in Continual Pre-training of Domain-Specific LLMs

EMNLP 2025

Data efficiency is crucial in domain-specific continual pre-training (CPT) of large language models (LLMs), especially under resource constraints. Aiming for “small data, big impact,” this work addresses the limitations of existing domain-specific data selection strategies, which often rely on scarc

2025

From Remembering to Metacognition: Do Existing Benchmarks Accurately Evaluate LLMs?

EMNLP 2025

Despite the rapid development of large language models (LLMs), existing benchmark datasets often focus on low-level cognitive tasks, such as factual recall and basic comprehension, while providing limited coverage of higher-level reasoning skills, including analysis, evaluation, and creation. In thi

Cited by 0SourcePDFScholar
2025

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

AAAI 2025technical

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale d…

2025

Single Domain Generalization for Few-Shot Counting via Universal Representation Matching

CVPR 2025poster

Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to unseen scenarios. This falls into the realm of single domain generalization that remains unexplored in few-shot countin…

2023

GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image

NeurIPS 2023poster

The extraordinary ability of generative models to generate photographic images has intensified concerns about the spread of disinformation, thereby leading to the demand for detectors capable of distinguishing between AI-generated fake images and real images. However, the lack of large datasets cont…

Cited by 137SourcePDFScholar
2022

How Well Does Self-Supervised Pre-Training Perform with Streaming Data?

ICLR 2022poster

Prior works on self-supervised pre-training focus on the joint training scenario, where massive unlabeled data are assumed to be given as input all at once, and only then is a learner trained. Unfortunately, such a problem setting is often impractical if not infeasible since many real-world tasks re…

Cited by 39SourcePDFScholar