← Search

Yongwei Nie

16 accepted papers

2026

Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability

ICLR 2026poster

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods, usually optimizing for high temporal Intersection-over-Union (IoU), frequently struggle to accurately recognize or und…

Cited by 0SourceScholar
2026

T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal dynamics via positional encoding, text-based timestamps, or visu

Cited by 0SourceScholar
2025

EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completion

CVPR 2025poster

This paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes on…

2025

Occlusion-Insensitive Talking Head Video Generation via Facelet Compensation

AAAI 2025technical

Talking head video generation involves animating a still face image using facial motion cues derived from a driving video to replicate target poses and expressions. Traditional methods often rely on the assumption that the relative positions of facial keypoints remain unchanged. However, this assump…

Cited by 0SourcePDFScholar
2025

RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation

ICLR 2025poster

Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises because existing methods struggle to maintain consistency across…

Cited by 3SourcePDFScholar
2025

Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection

NeurIPS 2025poster

3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in cap…

Cited by 0SourceScholar
2025

Simplification Is All You Need against Out-of-Distribution Overconfidence

CVPR 2025poster

Deep neural networks (DNNs) often exhibit out-of-distribution (OOD) overconfidence, producing overly confident predictions on OOD samples. We attribute this issue to the inherent over-complexity of DNNs and investigate two key aspects: capacity and nonlinearity. First, we demonstrate that reducing m…

Cited by 3SourcePDFScholar
2025

When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired Data

AAAI 2025technical

We present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previ…

Cited by 0SourcePDFScholar
2024

Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery

NeurIPS 2024poster

Human Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pret…

2024

Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection

ECCV 2024poster

"Video Anomaly Detection (VAD) has been extensively studied under the settings of One-Class Classification (OCC) and Weakly-Supervised learning (WS), which however both require laborious human-annotated normal/abnormal labels. In this paper, we study Unsupervised VAD (UVAD) that does not depend on a…

2024

Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses

ECCV 2024poster

"Besides a 3D mesh, Human Mesh Recovery (HMR) methods usually need to estimate a camera for computing 2D reprojection loss. Previous approaches may encounter the following problem: both the mesh and camera are not correct but the combination of them can yield a low reprojection loss. To alleviate th…

2023

Fine-Grained Face Swapping via Regional GAN Inversion

CVPR 2023poster

We present a novel paradigm for high-fidelity face swapping that faithfully preserves the desired subtle geometry and texture details. We rethink face swapping from the perspective of fine-grained face editing, i.e., editing for swapping (E4S), and propose a framework that is based on the explicit d…

2022

Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion Prediction

CVPR 2022poster

This paper presents a high-quality human motion prediction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good initial guess of the future poses is very helpful in improving the forecasting accuracy. This motivates us to propose…

Cited by 140PDFcodeScholar
2021

A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction

ICCV 2021poster

In this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns…

Cited by 278PDFcodeScholar
2021

MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction

ICCV 2021poster

Human motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract…

Cited by 262PDFcodeScholar