← Search

Cheng Xu

17 accepted papers

2026

ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection

ICLR 2026poster

Surgical triplet detection is a critical task in surgical video analysis, with significant implications for performance assessment and training novice surgeons. However, existing datasets like CholecT50 lack precise spatial bounding box annotations, rendering triplet classification at the image leve…

Cited by 0SourceScholar
2026

Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual Inputs

ICLR 2026poster

Vision-language models (VLMs) are increasingly replacing traditional OCR pipelines. However, they often hallucinate on lossy visual inputs, such as visually degraded document images, producing fluent yet incorrect text without signaling uncertainty. This occurs because current post-training emphasiz…

Cited by 0SourceScholar
2025

DCR: Quantifying Data Contamination in LLMs Evaluation

EMNLP 2025

The rapid advancement of large language models (LLMs) has heightened concerns about benchmark data contamination (BDC), where models inadvertently memorize evaluation data during the training process, inflating performance metrics, and undermining genuine generalization assessment. This paper introd

2025

Details Enhancement in Unsigned Distance Field Learning for High-fidelity 3D Surface Reconstruction

AAAI 2025technical

While Signed Distance Fields (SDF) are well-established for modeling watertight surfaces, Unsigned Distance Fields (UDF) broaden the scope to include open surfaces and models with complex inner structures. Despite their flexibility, UDFs encounter significant challenges in high-fidelity 3D reconstru…

Cited by 0SourcePDFScholar
2025

Dr. Tongue: Sign-Oriented Multi-label Detection for Remote Tongue Diagnosis

AAAI 2025technical

Tongue diagnosis is a vital tool in both Western and Traditional Chinese Medicine, providing key insights into a patient's health by analyzing tongue attributes. The COVID-19 pandemic has heightened the need for accurate remote medical assessments, emphasizing the importance of precise tongue attrib…

2025

FR²Seg: Continual Segmentation Across Multiple Sites via Fourier Style Replay and Adaptive Consistency Regularization

AAAI 2025technical

In clinical imaging, medical segmentation networks typically require continually adapting to new data from multiple sites over time, as aggregating all data for learning at once can be impractical due to storage limitations and privacy concerns. However, existing methods basically overlook domain-s…

2025

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

ACL 2025long

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded historical documents.First, we synthesize document-image pairs…

2025

SSA: Semantic Contamination of LLM-Driven Fake News Detection

EMNLP 2025

Benchmark data contamination (BDC) silently inflate the evaluation performance of large language models (LLMs), yet current work on BDC has centered on direct token overlap (data/label level), leaving the subtler and equally harmful semantic level BDC largely unexplored. This gap is critical in fake

2025

StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models

NeurIPS 2025poster

The advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods have made progress toward unified solutions, their reliance on post hoc processing…

Cited by 0SourceScholar
2025

TripleFact: Defending Data Contamination in the Evaluation of LLM-driven Fake News Detection

ACL 2025long

The proliferation of large language models (LLMs) has introduced unprecedented challenges in fake news detection due to benchmark data contamination (BDC), where evaluation benchmarks are inadvertently memorized during the pre-training, leading to the inflated performance metrics. Traditional evalua…

2024

Effective Synthetic Data and Test-Time Adaptation for OCR Correction

EMNLP 2024main

Post-OCR technology is used to correct errors in the text produced by OCR systems. This study introduces a method for constructing post-OCR synthetic data with different noise levels using weak supervision. We define Character Error Rate (CER) thresholds for “effective” and “ineffective” synthetic d…

Cited by 1SourcePDFScholar
2024

Reinforcement Learning Compensated Filter for Multi-Agents Cooperative Localization

ICASSP 2024accepted

Accurate and real-time location tracking is vital for various applications in public safety and the military, particularly in search and rescue missions. Traditional filtering localization algorithms are more effective in linear environments and require precise initial estimates and system noise for…

Cited by 0SourceScholar
2023

CIRI: Curricular Inactivation for Residue-aware One-shot Video Inpainting

ICCV 2023poster

Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated f…

Cited by 9PDFcodeScholar
2022

Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification

ICASSP 2022accepted

Multi-label video classification (MLVC) is a long-standing and challenging research problem in video signal analysis. Generally, there exist many complex action labels in real-world videos and these actions are with inherent dependencies at both spatial and temporal domains. Motivated by this observ…

Cited by 0SourceScholar
2022

W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization

ICASSP 2022accepted

Weakly-supervised temporal action localization (WTAL) is a long-standing and challenging research problem in video signal analysis. It is to localize the action segments in the video given only video-level labels. The key to this task is understanding how the diverse actions interact. In this paper,…

Cited by 0SourceScholar
2019

A Novel Fractional Order Derivate Based Log-demons with Driving Force for High Accurate Image Registration

ICASSP 2019accepted

Image registration methods based on Thirion's demons method update displacement field by the image gradient obtained by integer order derivate. However, the fractional order derivate is superior to integral order derivate for computing image gradient under weak texture or smooth regions. To obtain h…

Cited by 0SourceScholar
2019

Enhancing 2D Representation via Adjacent Views for 3D Shape Retrieval

ICCV 2019poster

Multi-view shape descriptors obtained from various 2D images are commonly adopted in 3D shape retrieval. One major challenge is that significant shape information are discarded during 2D view rendering through projection. In this paper, we propose a convolutional neural network based method, CenterN…

Cited by 23PDFScholar