← Search

Feng Xu

43 accepted papers

2026

CofactGVR: Counterfactual Intervention for Grounded Visual Reasoning

ICML 2026poster

Despite rapid progress in Grounded Visual Reasoning (GVR) with MLLMs and RL-style fine-tuning, existing approaches often lack effective learning signals for intermediate grounding decisions and are prone to shortcut solutions. In this work, we explicitly decompose GVR into Evidence Generation follow…

Cited by 0SourceScholar
2026

Fair Conformal Classification via Learning Representation-Based Groups

ICLR 2026poster

Conformal prediction methods provide statistically rigorous marginal coverage guarantees for machine learning models, but such guarantees fail to account for algorithmic biases, thereby undermining fairness and trust. This paper introduces a fair conformal inference framework for classification task…

Cited by 0SourceScholar
2026

Learnable Data Augmentation and Contrastive Pre-training for Temporal Link Prediction

IJCAI 2026

Link prediction is a foundational task in temporal graphs. While temporal graph neural networks exhibit commendable performance, they are often criticized for providing inadequate representations, especially under limited data. Contrastive learning has been introduced as a solution for graph pre-tra

Cited by 0Scholar
2026

PEERING INTO THE UNKNOWN: ACTIVE VIEW SELECTION WITH NEURAL UNCERTAINTY MAPS FOR 3D RECONSTRUCTION

ICLR 2026poster

Imagine trying to understand the shape of a teapot by viewing it from the front—you might see the spout, but completely miss the handle. Some perspectives naturally provide more information than others. How can an AI system determine which viewpoint offers the most valuable insight for accurate and…

Cited by 0SourcecodeScholar
2026

VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the Wild

AAAI 2026technical

Reconstructing topologically consistent facial geometry is crucial for the digital avatar creation pipelines. Existing methods either require tedious manual efforts, lack generalization to in-the-wild data, or are constrained by the limited expressiveness of 3D Morphable Models. To address these lim

Cited by 0SourcePDFScholar
2026

WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Rendering

CVPR 2026

Existing methods achieve high-quality facial albedo capture under controllable lighting, which increases capture cost and limits usability. We propose WildCap, a novel method for high-quality facial albedo capture from a smartphone video recorded in the wild. To disentangle high-quality albedo from

Cited by 0SourcecodeScholar
2025

A spectrum-enhanced attention model for semantic segmentation of remote sensing images

ICASSP 2025accepted

Semantic segmentation of remote sensing images (RSIs) is essential for applications such as environmental monitoring, urban planning, and disaster management. Convolutional Neural Networks (CNNs) and their variants struggle to capture comprehensive spectral context for learning discriminative repres…

Cited by 0SourceScholar
2025

Consistency-aware Self-Training for Iterative-based Stereo Matching

CVPR 2025poster

Iterative-based methods have become mainstream in stereo matching due to their high performance. However, these methods heavily rely on labeled data and face challenges with unlabeled real-world data. To this end, we propose a consistency-aware self-training framework for iterative-based stereo matc…

Cited by 0SourcePDFScholar
2025

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

ICCV 2025poster

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for repres…

2025

MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic Disturbances

ICCV 2025poster

This paper proposes a novel method, named MagShield, designed to address the issue of magnetic disturbances in sparse inertial motion capture (MoCap) systems. Existing Inertial Measurement Units (IMUs) are prone to orientation estimation errors in magnetically disturbed environments, limiting the pr…

2025

Multiple Sclerosis Detection with Reinforcement Learning and Differential Evolution

ICASSP 2025accepted

Multiple Sclerosis (MS) disrupts nerve communication, potentially leading to permanent damage. Convolutional Neural Networks (CNNs) are commonly recommended to accelerate magnetic resonance imaging (MRI) analysis for MS. Traditional CNN-based methods often face challenges with feature selection, imb…

Cited by 0SourceScholar
2025

Simulate, Refine and Integrate: Strategy Synthesis for Efficient SMT Solving

IJCAI 2025

Satisfiability Modulo Theories (SMT) solvers are crucial in many applications, yet their performance is often a bottleneck. This paper introduces SIRISMT, a novel framework that employs machine learning techniques for the automatic synthesis of efficient SMT-solving strategies. Specifically, SIRISMT

2025

Teeth Reconstruction and Performance Capture Using a Phone Camera

ICCV 2025poster

We present the first method for personalized dental shape reconstruction and teeth-inclusive facial performance capture using only a single phone camera. Our approach democratizes high-quality facial avatars through a non-invasive, low-cost setup by addressing the ill-posed monocular capture problem…

2024

Demonstrating Event-Triggered Investigation and Sample Collection for Human Scientists using Field Robots and Large Foundation Models

RSS 2024poster

In this paper, we introduce a pioneering end-to-end system demonstrated on a team of robots and sensors, designed to augment scientific exploration and discovery for human scientists in remote or inaccessible environments. We demonstrate and analyse our system's capability in a mock-up test-bed scen…

2024

Inspecting Prediction Confidence for Detecting Black-Box Backdoor Attacks

AAAI 2024technical

Backdoor attacks have been shown to be a serious security threat against deep learning models, and various defenses have been proposed to detect whether a model is backdoored or not. However, as indicated by a recent black-box attack, existing defenses can be easily bypassed by implanting the backdo…

Cited by 10SourcePDFScholar
2024

Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing Images

ICASSP 2024accepted

Transformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However,…

Cited by 0SourceScholar
2024

Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear Jacket

CVPR 2024poster

Existing wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing limiting their application in everyday life. In this paper we introduce Loose Inertial Poser a novel motion capture solution with high wearing comfortableness by integrating fou…

2024

OptiState: State Estimation of Legged Robots using Gated Networks with Transformer-based Vision and Kalman Filtering

ICRA 2024poster

State estimation for legged robots is challenging due to their highly dynamic motion and limitations imposed by sensor accuracy. By integrating Kalman filtering, optimization, and learning-based modalities, we propose a hybrid solution that combines proprioception and exteroceptive information for e…

Cited by 6SourcecodeScholar
2024

Relightable and Animatable Neural Avatars from Videos

AAAI 2024technical

Lightweight creation of 3D digital avatars is a highly desirable but challenging task. With only sparse videos of a person under unknown illumination, we propose a method to create relightable and animatable neural avatars, which can be used to synthesize photorealistic images of humans under novel…

2023

EditableNeRF: Editing Topologically Varying Neural Radiance Fields by Key Points

CVPR 2023poster

Neural radiance fields (NeRF) achieve highly photo-realistic novel-view synthesis, but it's a challenging problem to edit the scenes modeled by NeRF-based methods, especially for dynamic scenes. We propose editable neural radiance fields that enable end-users to easily edit dynamic scenes and even s…

2023

Transmit Energy Focusing For Parameter Estimation in Transmit Beamspace Slow-Time MIMO Radar

ICASSP 2023accepted

Recently, Parallel Factor-Direct (PARAFAC-Direct) method has been proposed for parameter estimation including velocity disambiguation for Doppler Division Multiple Access (DDMA) Multiple-Input Multiple-Output (MIMO) radar. However, DDMA MIMO radar spreads the overall transmit energy into the entire…

Cited by 0SourceScholar
2023

UniCOQE: Unified Comparative Opinion Quintuple Extraction As A Set

ACL 2023findings

Comparative Opinion Quintuple Extraction (COQE) aims to identify comparative opinion sentences in product reviews, extract comparative opinion elements in the sentences, and then incorporate them into quintuples. Existing methods decompose the COQE task into multiple primary subtasks and then solve…

2022

An Invisible Black-Box Backdoor Attack through Frequency Domain

ECCV 2022poster

"Backdoor attacks have been shown to be a serious threat against deep learning systems such as biometric authentication and autonomous driving. An effective backdoor attack could enforce the model misbehave under certain predefined conditions, i.e., triggers, but behave normally otherwise. The trigg…

2022

OcclusionFusion: Occlusion-Aware Motion Estimation for Real-Time Dynamic 3D Reconstruction

CVPR 2022poster

RGBD-based real-time dynamic 3D reconstruction suffers from inaccurate inter-frame motion estimation as errors may accumulate with online tracking. This problem is even more severe for single-view-based systems due to strong occlusions. Based on these observations, we propose OcclusionFusion, a nove…

Cited by 40PDFcodeScholar
2022

Physical Inertial Poser (PIP): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors

CVPR 2022poster

Motion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global posit…

Cited by 200PDFScholar
2022

Structure-Aware Editable Morphable Model for 3D Facial Detail Animation and Manipulation

ECCV 2022poster

"Morphable models are essential for the statistical modeling of 3D faces. Previous works on morphable models mostly focus on large-scale facial geometry but ignore facial details. This paper augments morphable models in representing facial details by learning a Structure-aware Editable Morphable Mod…

2021

Constrained Tensor Decomposition for 2d DOA Estimation In Transmit Beamspace Mimo Radar with Subarrays

ICASSP 2021accepted

In this paper, a constrained tensor decomposition method that enables two dimensional (2D) direction of arrival (DOA) estimation for transmit beamspace (TB) Multiple-Input Multiple-Output (MIMO) radar with subarrays is proposed. Specifically, a higher-order tensor model is designed to collect the re…

Cited by 0SourceScholar
2021

Cross-modal Domain Adaptation for Cost-Efficient Visual Reinforcement Learning

NeurIPS 2021poster

In visual-input sim-to-real scenarios, to overcome the reality gap between images rendered in simulators and those from the real world, domain adaptation, i.e., learning an aligned representation space between simulators and the real world, then training and deploying policies in the aligned represe…

2021

Monocular Real-Time Full Body Capture With Inter-Part Correlations

CVPR 2021poster

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits correlations between body and hands at high computational e…

Cited by 72PDFScholar
2021

Regret Minimization Experience Replay in Off-Policy Reinforcement Learning

NeurIPS 2021poster

In reinforcement learning, experience replay stores past samples for further reuse. Prioritized sampling is a promising technique to better utilize these samples. Previous criteria of prioritization include TD error, recentness and corrective feedback, which are mostly heuristically designed. In thi…

2020

Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal Data

CVPR 2020poster

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image da…

Cited by 257PDFcodeScholar
2020

SQUIRL: Robust and Efficient Learning from Video Demonstration of Long-Horizon Robotic Manipulation Tasks

IROS 2020poster

Recent advances in deep reinforcement learning (RL) have demonstrated its potential to learn complex robotic manipulation tasks. However, RL still requires the robot to collect a large amount of real-world experience. To address this problem, recent works have proposed learning from expert demonstra…

Cited by 22SourceScholar
2020

Trading Personalization for Accuracy: Data Debugging in Collaborative Filtering

NeurIPS 2020poster

Collaborative filtering has been widely used in recommender systems. Existing work has primarily focused on improving the prediction accuracy mainly via either building refined models or incorporating additional side information, yet has largely ignored the inherent distribution of the input rating…

2018

DDRNet: Depth Map Denoising and Refinement for Consumer Depth Cameras Using Cascaded CNNs

ECCV 2018poster

Consumer depth sensors are more and more popular and come to our daily lives marked by its recent integration in the latest Iphone X. However, they still suffer from heavy noises which limit their applications. Although plenty of progresses have been made to reduce the noises and boost geometric det…

2017

BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera

ICCV 2017poster

We propose BodyFusion, a novel real-time geometry fusion method that can track and reconstruct non-rigid surface motion of a human performance using a single consumer-grade depth camera. To reduce the ambiguities of the non-rigid deformation parameterization on the surface graph nodes, we take advan…

Cited by 200PDFScholar
2017

Decoder Network Over Lightweight Reconstructed Feature for Fast Semantic Style Transfer

ICCV 2017poster

Recently, the community of style transfer is trying to incorporate semantic information into traditional system. This practice achieves better perceptual results by transferring the style between semantically-corresponding regions. Yet, few efforts are invested to address the computation bottleneck…

Cited by 71PDFScholar
2015

Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 Regularization

ICCV 2015poster

We present a new motion tracking method to robustly reconstruct non-rigid geometries and motions from single view depth inputs captured by a consumer depth sensor. The idea comes from the observation of the existence of intrinsic articulated subspace in most of non-rigid motions. To take advantage o…

Cited by 146PDFScholar