← Search

Xuan Feng

5 accepted papers

2026

SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes

CVPR 2026

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on differentiable renderings without any flow supervision. Besi

Cited by 0SourcecodeScholar
2025

Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction

AAAI 2025technical

Unnatural text correction aims to automatically detect and correct spelling errors or adversarial perturbation errors in sentences. Existing methods typically rely on fine-tuning or adversarial training to correct errors, which have achieved significant success. However, these methods exhibit poor g…

2025

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

ICML 2025poster

Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that…

Cited by 0SourcePDFScholar
2022

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

IJCAI 2022poster

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose estimation methods focus on adults, lacking publicly benchmark for infant pose esti…