← Search

Shaofei Wang

14 accepted papers

2026

Collaborative Enhancement of Large and Small Models for Question Answering via Dual Knowledge Transfer

AAAI 2026technical

Our statistical analysis reveals a complementary phenomenon between large language model-based question answering (QA) and small model-based QA. To facilitate dual knowledge transfer between these two paradigms, this paper introduces a collaborative enhancement method of large and small models for q

Cited by 0SourcePDFScholar
2026

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

CVPR 2026

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark

Cited by 0SourceScholar
2026

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

CVPR 2026

Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-

Cited by 0SourcecodeScholar
2025

Logic-Thinker: Teaching Large Language Models to Think more Logically.

EMNLP 2025

Recent Large Reasoning Models (LRMs) have demonstrated the ability to generate long chains of thought (LongCoT) before arriving at a final conclusion. Despite remarkable breakthroughs in complex reasoning capabilities, LongCoT still faces challenges such as redundancy and logical incoherence. To add

Cited by 0SourcePDFScholar
2024

3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

CVPR 2024poster

We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) achieve high-quality novel-view/novel-pose image synthesis but often require days of training and are extremely slow at in…

Cited by 123SourcePDFScholar
2024

IntrinsicAvatar: Physically Based Inverse Rendering of Dynamic Humans from Monocular Videos via Explicit Ray Tracing

CVPR 2024poster

We present IntrinsicAvatar a novel approach to recovering the intrinsic properties of clothed human avatars including geometry albedo material and environment lighting from only monocular videos. Recent advancements in human-based neural rendering have enabled high-quality geometry and appearance re…

Cited by 12SourcePDFScholar
2024

Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation

CVPR 2024poster

Recent advances in generative diffusion models have enabled the previously unfeasible capability of generating 3D assets from a single input image or a text prompt. In this work we aim to enhance the quality and functionality of these models for the task of creating controllable photorealistic human…

2023

Synthesizing Diverse Human Motions in 3D Indoor Scenes

ICCV 2023poster

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on high-quality training sequences that contain captured human motions and the 3D scenes they interact with. Ho…

Cited by 66PDFcodeScholar
2022

ARAH: Animatable Volume Rendering of Articulated Human SDFs

ECCV 2022poster

"Combining human body models with differentiable rendering has recently enabled animatable avatars of clothed humans from sparse sets of multi-view RGB videos. While state-of-the-art approaches achieve a realistic appearance with neural radiance fields (NeRF), the inferred geometry often lacks detai…

Cited by 150SourcePDFScholar
2022

Compositional Human-Scene Interaction Synthesis with Semantic Control

ECCV 2022poster

"Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Recent methods mainly focus on modeling geometric relations between 3D environments and humans, where the high-level semantics of t…

2021

MetaAvatar: Learning Animatable Clothed Human Models from Few Depth Images

NeurIPS 2021poster

In this paper, we aim to create generalizable and controllable neural signed distance fields (SDFs) that represent clothed humans from monocular depth observations. Recent advances in deep learning, especially neural implicit representations, have enabled human shape reconstruction and controllable…

2018

Accelerating Dynamic Programs via Nested Benders Decomposition with Application to Multi-Person Pose Estimation

ECCV 2018poster

We present a novel approach to solve dynamic programs (DP), which are frequent in computer vision, on tree-structured graphs with exponential node state space. Typical DP approaches have to enumerate the joint state space of two adjacent nodes on every edge of the tree to compute the optimal message…

Cited by 14SourcePDFScholar
2017

Tracking Objects with Higher Order Interactions via Delayed Column Generation

AISTATS 2017poster

We study the problem of multi-target tracking and data association in video. We formulate this in terms of selecting a subset of high-quality tracks subject to the constraint that no pair of selected tracks is associated with a common detection (of an object). This objective is equivalent to the cla…

Cited by 13SourcePDFScholar