← Search

Ruotao Xu

8 accepted papers

2025

A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse Artifacts

CVPR 2025poster

Structured artifacts are semi-regular, repetitive patterns that closely intertwine with genuine image content, making their removal highly challenging. In this paper, we introduce the Scale-Adaptive Deformable Transformer, an network architecture specifically designed to eliminate such artifacts fro…

2025

Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video Restoration

AAAI 2025technical

Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete P…

2025

RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting

AAAI 2025technical

Face retouching aims to remove facial imperfections from image and videos while at the same time preserving face attributes. The existing methods are designed to perform non-interactive end-to-end retouching, while the ability to interact with users is highly demanded in downstream applications. In…

Cited by 0SourcePDFScholar
2025

Self-Correcting Robot Manipulation via Gaussian-Splatted Foresight

AAAI 2025technical

Language-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the…

Cited by 0SourcePDFScholar
2025

Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding

CVPR 2025poster

The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual re…

Cited by 0SourcePDFScholar