← Search

Wentao Wang

20 accepted papers

2026

EgoTwin: Dreaming Body and View in First Person

ICLR 2026poster

While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the wearer's body movements. To bridge this gap, we introduce a novel task of joint…

Cited by 0SourceScholar
2026

GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space

ICLR 2026poster

In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key challenge lies in handling heterogeneous features from agents equipped with different sensing modalities or model architectures, which complicates data fusion.…

Cited by 0SourcecodeScholar
2026

High-Quality Full-Head 3D Avatar Generation from Any Single Portrait Image

AAAI 2026technical

In this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modelin

Cited by 0SourcePDFScholar
2026

Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning

AAAI 2026technical

The growing exploration of Large Language Models (LLM) and Vision-Language Models (VLM) has opened avenues for enhancing the effectiveness of reinforcement learning (RL). However, existing LLM-based RL methods often focus on the guidance of control policy and encounter the challenge of limited repre

Cited by 0SourcePDFScholar
2025

BEVSync: Asynchronous Data Alignment for Camera-based Vehicle-Infrastructure Cooperative Perception Under Uncertain Delays

AAAI 2025technical

Vehicle-to-infrastructure (V2I) cooperative perception systems can enhance the sensing abilities of autonomous vehicles. Existing V2I solutions often consider LiDARs devices instead of cameras, the most prevalent sensors with low cost and wide installation. In addition, a major challenge that has be…

Cited by 0SourcePDFScholar
2025

GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data

NeurIPS 2025poster

Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities i…

Cited by 0SourceScholar
2025

High-Fidelity Lightweight Mesh Reconstruction from Point Clouds

CVPR 2025highlight

Recently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To…

Cited by 0SourcePDFScholar
2025

TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection

ICASSP 2025accepted

Object detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applicatio…

Cited by 0SourceScholar
2024

CosmicMan: A Text-to-Image Foundation Model for Humans

CVPR 2024highlight

We present CosmicMan a text-to-image foundation model specialized for generating high-fidelity human images. Unlike current general-purpose foundation models that are stuck in the dilemma of inferior quality and text-image misalignment for humans CosmicMan enables generating photo-realistic human im…

2024

InterCoop: Spatio-Temporal Interaction Aware Cooperative Perception for Networked Vehicles

ICRA 2024poster

In autonomous driving, cooperative perception through vehicle-to-vehicle (V2V) communication is considered crucial for enhancing traffic safety and efficiency. However, existing methods often simplify the handling of perception data from multiple vehicles. In these approaches, the egovehicle aggrega…

Cited by 2SourceScholar
2023

Efficient Planar Pose Estimation via UWB Measurements

ICRA 2023poster

State estimation is an essential part of autonomous systems. Integrating the Ultra-Wideband (UWB) technique has been shown to correct the long-term estimation drift and bypass the complexity of loop closure detection. However, few works on robotics treat UWB as a stand-alone state estimation solutio…

Cited by 15SourcecodeScholar
2023

The KFIoU Loss for Rotated Object Detection

ICLR 2023poster

Differing from the well-developed horizontal object detection area whereby the computing-friendly IoU based loss is readily adopted and well fits with the detection metrics, rotation detectors often involve a more complicated loss based on SkewIoU which is unfriendly to gradient-based training. In t…

Cited by 239SourcePDFScholar
2021

Dense Label Encoding for Boundary Discontinuity Free Rotation Detection

CVPR 2021poster

Rotation detection serves as a fundamental building block in many visual applications involving aerial image, scene text, and face etc. Differing from the dominant regression-based approaches for orientation estimation, this paper explores a relatively less-studied methodology based on classificatio…

Cited by 350PDFScholar
2021

Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence

NeurIPS 2021poster

Existing rotated object detectors are mostly inherited from the horizontal detection paradigm, as the latter has evolved into a well-developed area. However, these detectors are difficult to perform prominently in high-precision detection due to the limitation of current regression loss design, espe…

2021

Parallel Multi-Resolution Fusion Network for Image Inpainting

ICCV 2021poster

Conventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow laye…

Cited by 39PDFScholar
2021

Rethinking Rotated Object Detection with Gaussian Wasserstein Distance Loss

ICML 2021spotlight

Boundary discontinuity and its inconsistency to the final detection metric have been the bottleneck for rotating detection regression loss design. In this paper, we propose a novel regression loss based on Gaussian Wasserstein distance as a fundamental approach to solve the problem. Specifically, th…