← Search

Xiangyu Xu

33 accepted papers

2026

Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs

AAAI 2026technical

Deep Neural Networks (DNNs), as valuable intellectual property, face unauthorized use. Existing protections, such as digital watermarking, are largely passive; they provide only post-hoc ownership verification and cannot actively prevent the illicit use of a stolen model. This work proposes a proact

Cited by 0SourcePDFScholar
2026

Diffusion-Based mmWave Radar Point Cloud Enhancement Driven by Range Images

RA-L 2026

Millimeter-wave (mmWave) radar has attracted significant attention in robotics and autonomous driving due to its robustness in harsh environments. However, the radar point clouds are typically sparse and noisy, which limits its futher development. Traditional mmWave radar enhancement approaches ofte

Cited by 6SourceScholar
2026

EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI

AAAI 2026technical

Generative modeling has recently shown remarkable promise for visuomotor policy learning, enabling flexible and expressive control across diverse embodied AI tasks. However, existing generative policies often struggle with data inefficiency, requiring large-scale demonstrations, and sampling ineffic

Cited by 0SourcePDFScholar
2026

MIMO-LP: A Multi-Input Multi-Output Framework for Subgraph-based Link Prediction

ICML 2026poster

Link prediction (LP) is a fundamental problem in graph learning and can be broadly categorized into node-based and subgraph-based approaches. While subgraph-based LP methods often achieve superior predictive performance by exploiting localized structural information, they suffer from efficiency bott…

Cited by 0SourceScholar
2025

ActiveGAMER: Active GAussian Mapping through Efficient Rendering

CVPR 2025poster

We introduce ActiveGAMER, an active mapping system that utilizes 3D Gaussian Splatting (3DGS) to achieve high-quality, real-time scene mapping and exploration. Unlike traditional NeRF-based methods, which are computationally demanding and restrict active mapping performance, our approach leverages t…

Cited by 2SourcePDFScholar
2025

Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoning

EMNLP 2025

Introducing **MARK**, the **M**ulti-st**A**ge **R**easoning framewor**K** for cultural value survey response simulation, designed to enhance the accuracy, steerability, and interpretability of large language models in this task. The system is inspired by the type dynamics theory in the MBTI psycholo

Cited by 0SourcePDFScholar
2025

EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients

ICCV 2025poster

Federated learning research has recently shifted from Convolutional Neural Networks (CNNs) to Vision Transformers (ViTs) due to their superior capacity. ViTs training demands higher computational resources due to the lack of 2D inductive biases inherent in CNNs. However, efficient federated training…

Cited by 0SourcePDFScholar
2025

GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models

ICLR 2025poster

In this paper, we introduce GoodDrag, a novel approach to improve the stability and image quality of drag editing. Unlike existing methods that struggle with accumulated perturbations and often result in distortions, GoodDrag introduces an AlDD framework that alternates between drag and denoising op…

2025

PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance Fields

ICRA 2025

Identifying spatially complete planar primitives from visual data is a crucial task in computer vision. Prior methods are largely restricted to either 2D segment recovery or simplifying 3D structures, even with extensive plane annotations. We present PlanarNeRF, a novel framework capable of detectin

Cited by 8SourceScholar
2025

Robust Neural Rendering in the Wild with Asymmetric Dual 3D Gaussian Splatting

NeurIPS 2025spotlight

3D reconstruction from in-the-wild images remains a challenging task due to inconsistent lighting conditions and transient distractors. Existing methods typically rely on heuristic strategies to handle the low-quality training data, which often struggle to produce stable and consistent reconstructio…

Cited by 0SourceScholar
2025

Tailless Flapping-Wing Robot With Bio-Inspired Elastic Passive Legs for Multi-Modal Locomotion

RA-L 2025

Flapping-wing robots offer significant versatility; however, achieving efficient multi-modal locomotion remains challenging. This paper presents the design, modeling, and experimentation of a novel tailless flapping-wing robot with three independently actuated pairs of wings. Inspired by the leg mor

Cited by 4SourceScholar
2025

WMarkGPT: Watermarked Image Understanding via Multimodal Large Language Models

ICML 2025poster

Invisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in…

2024

Motion-adaptive Separable Collaborative Filters for Blind Motion Deblurring

CVPR 2024poster

Eliminating image blur produced by various kinds of motion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in feature space. These practices not only prevent the capture of spatially variable mot…

2024

NARUTO: Neural Active Reconstruction from Uncertain Target Observations

CVPR 2024poster

We present NARUTO a neural active reconstruction system that combines a hybrid neural representation with uncertainty learning enabling high-fidelity surface reconstruction. Our approach leverages a multi-resolution hash-grid as the mapping backbone chosen for its exceptional convergence speed and c…

2023

NU-MCC: Multiview Compressive Coding with Neighborhood Decoder and Repulsive UDF

NeurIPS 2023poster

Remarkable progress has been made in 3D reconstruction from single-view RGB-D inputs. MCC is the current state-of-the-art method in this field, which achieves unprecedented success by combining vision Transformers with large-scale training. However, we identified two key limitations of MCC: 1) The T…

2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

ICCV 2023poster

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Seco…

Cited by 24PDFScholar
2022

BasicVSR++: Improving Video Super-Resolution With Enhanced Propagation and Alignment

CVPR 2022poster

A recurrent structure is a popular framework choice for the task of video super-resolution. The state-of-the-art method BasicVSR adopts bidirectional propagation with feature alignment to effectively exploit information from the entire input video. In this study, we redesign BasicVSR by proposing se…

Cited by 541PDFcodeScholar
2022

Geometry-Guided Progressive NeRF for Generalizable and Efficient Neural Human Rendering

ECCV 2022poster

"In this work we develop a generalizable and efficient Neural Radiance Field (NeRF) pipeline for high-fidelity free-viewpoint human body synthesis under settings with sparse camera views. Though existing NeRF-based methods can synthesize rather realistic details for human body, they tend to produce…

Cited by 49SourcePDFScholar
2022

Investigating Tradeoffs in Real-World Video Super-Resolution

CVPR 2022poster

The diversity and complexity of degradations in real-world video super-resolution (VSR) pose non-trivial challenges in inference and training. First, while long-term propagation leads to improved performance in cases of mild degradations, severe in-the-wild degradations could be exaggerated through…

Cited by 125PDFcodeScholar
2021

GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution

CVPR 2021poster

We show that pre-trained Generative Adversarial Networks (GANs), e.g., StyleGAN, can be used as a latent bank to improve the restoration quality of large-factor image super-resolution (SR). While most existing SR approaches attempt to generate realistic textures through learning with adversarial los…

Cited by 315PDFcodeScholar
2020

3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning

ECCV 2020poster

3D human shape and pose estimation from monocular images has been an active area of research in computer vision, having a substantial impact on the development of new applications, from activity recognition to creating virtual avatars. Existing deep learning methods for 3D human shape and pose estim…

2020

A Lightweight and Accurate Localization Algorithm Using Multiple Inertial Measurement Units

RA-L 2020

This paper proposes a novel inertial-aided localization approach by fusing information from multiple inertial measurement units (IMUs) and exteroceptive sensors. IMU is a low-cost motion sensor which provides measurements on angular velocity and gravity compensated linear acceleration of a moving pl

Cited by 64SourceScholar
2018

Monocular Depth Estimation with Affinity, Vertical Pooling, and Label Enhancement

ECCV 2018poster

While significant progress has been made in monocular depth estimation with Convolutional Neural Networks (CNNs) extracting absolute features, such as edges and textures, the depth constraint of neighboring pixels, namely relative features, has been mostly ignored by recent methods. To overcome this…

Cited by 148SourcePDFScholar
2018

Rendering Portraitures from Monocular Camera and Beyond

ECCV 2018poster

Shallow Depth-of-Field (DoF) is a desirable effect in photography which renders artistic photos. Usually, it requires single-lens reflex cameras and certain photography skills to generate such effects. Recently, dual-lens on cellphones is used to estimate scene depth and simulate DoF effects for por…

Cited by 32SourcePDFScholar
2017

Learning to Super-Resolve Blurry Face and Text Images

ICCV 2017poster

We present an algorithm to directly restore a clear high-resolution image from a blurry low-resolution input. This problem is highly ill-posed and the basic assumptions for existing super-resolution methods (requiring clear input) and deblurring methods (requiring high-resolution input) no longer ho…

Cited by 279PDFScholar