← Search

Rong Xie

16 accepted papers

2026

DP-DEGAUSS: DYNAMIC PROBABILISTIC GAUSSIAN DECOMPOSITION FOR EGOCENTRIC 4D SCENE RECONSTRUCTION

ICASSP 2026poster

Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic first-person scenes is challenging due to complex ego-motion, occlusions, and hand-object interactions. Existing decomposition methods are ill-suited,…

Cited by 0SourcePDFScholar
2025

H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting

NeurIPS 2025poster

Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space…

Cited by 0SourceScholar
2024

Cognitive Virtual Sensing Technique for Feedforward Active Noise Control

ICASSP 2024accepted

The virtual sensing (VS) technique enables an active noise control (ANC) system to estimate the virtual error signal for control using remote monitoring microphones. However, instances where noise characteristics and primary paths exhibit variations lead to a noticeable decline in performance for th…

Cited by 0SourceScholar
2024

Depth-Guided Robust and Fast Point Cloud Fusion NeRF for Sparse Input Views

AAAI 2024technical

Novel-view synthesis with sparse input views is important for real-world applications like AR/VR and autonomous driving. Recent methods have integrated depth information into NeRFs for sparse input synthesis, leveraging depth prior for geometric and spatial understanding. However, most existing work…

Cited by 6SourcePDFScholar
2024

Disentangled Clothed Avatar Generation from Text Descriptions

ECCV 2024poster

"In this paper, we introduce a novel text-to-avatar generation method that separately generates the human body and the clothes and allows high-quality animation on the generated avatar. While recent advancements in text-to-avatar generation have yielded diverse human avatars from text prompts, these…

Cited by 24SourcePDFScholar
2024

Hdrtvformer: Efficient Sdrtv-to-Hdrtv via Affine Transformation and Spatial-Aware Transformer

ICASSP 2024accepted

Recent works on reconstructing HDR videos in display format (HDRTV) suffer from high computational and memory requirements because they learn the SDRTV-to-HDRTV mapping directly in 4K resolution. This paper proposes an efficient SDRTV-to-HDRTV model (HDRTVFormer) that decomposes the HDRTV restoratio…

Cited by 0SourceScholar
2023

Boosting Video Object Segmentation via Space-Time Correspondence Learning

CVPR 2023poster

Current top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from…

2023

Divide and Conquer: a Two-Step Method for High Quality Face De-identification with Model Explainability

ICCV 2023poster

Face de-identification involves concealing the true identity of a face while retaining other facial characteristics. Current target-generic methods typically disentangle identity features in the latent space, using adversarial training to balance privacy and utility. However, this pattern often lead…

Cited by 22PDFcodeScholar
2022

A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution

ECCV 2022poster

"Online processing of compressed videos to increase their resolutions attracts increasing and broad attention. Video Super-Resolution (VSR) using recurrent neural network architecture is a promising solution due to its efficient modeling of long-range temporal dependencies. However, state-of-the-art…

Cited by 9SourcePDFScholar
2022

PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer towards Video Object Detection

ECCV 2022poster

"Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features at one stroke to enhance the feature. These methods, however, usually lack spatial information from neighboring frames a…

2021

Personalized and Invertible Face De-Identification by Disentangled Identity Information Manipulation

ICCV 2021poster

The popularization of intelligent devices including smartphones and surveillance cameras results in more serious privacy issues. De-identification is regarded as an effective tool for visual privacy protection with the process of concealing or replacing identity information. Most of the existing de-…

Cited by 80PDFScholar
2021

Region-Aware Adaptive Instance Normalization for Image Harmonization

CVPR 2021poster

Image composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground to be compatible with the background. Existing deep learning methods for harmonizing composite images directly learn an…

Cited by 144PDFcodeScholar
2020

Toward Fine-grained Facial Expression Manipulation

ECCV 2020poster

Facial expression manipulation aims at editing facial expression with a given condition. Previous methods edit an input image under the guidance of a discrete emotion label or absolute condition (e.g., facial action units) to possess the desired expression. However, these methods either suffer from…

2019

Selective Virtual Sensing Technique for Multi-channel Feedforward Active Noise Control Systems

ICASSP 2019accepted

The virtual sensing technique allows the active noise control (ANC) system to work with error microphones that are placed far from the desired zone of quietness (ZoQ). Conventionally, a training stage is required to obtain the auxiliary filters with the temporary error microphones placed in the ZoQ.…

Cited by 0SourceScholar
2018

Learning an Inverse Tone Mapping Network with a Generative Adversarial Regularizer

ICASSP 2018accepted

Transferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inver…

Cited by 0SourceScholar