← Search

Jingyi Yu

100 accepted papers

2026

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

ICML 2026poster

Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often limited to static predictions or require computationally intensive heuristic searches. We present CryoACE, an end-to-end f…

Cited by 0SourceScholar
2026

CylinderSplat: 3D Gaussian Splatting with Cylindrical Triplanes for Panoramic Novel View Synthesis

ICLR 2026poster

Feed-forward 3D Gaussian Splatting (3DGS) has shown great promise for real-time novel view synthesis, but its application to panoramic imagery remains challenging. Existing methods often rely on multi-view cost volumes for geometric refinement, which struggle to resolve occlusions in sparse-view sce…

Cited by 0SourcecodeScholar
2026

Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

ICML 2026poster

Diffusion Bridge and Flow Matching have both demonstrated compelling empirical performance in transformation between arbitrary distributions. However, there remains confusion about which approach is generally preferable, and the substantial discrepancies in their modeling assumptions and practical i…

Cited by 0SourceScholar
2026

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

ICML 2026poster

Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods struggle with effective multimodal modeling. They often rely on brute-force fusion strategies that ignore the structural dispa…

Cited by 0SourceScholar
2026

Improving 2D Diffusion Models for 3D Medical Imaging with Inter‑Slice Consistent Stochasticity

ICLR 2026poster

3D medical imaging is in high demand and essential for clinical diagnosis and scientific research. Currently, diffusion models have become an effective tool for medical imaging reconstruction thanks to their ability to learn rich, high‑quality data priors. However, learning the 3D data distribution…

Cited by 0SourcecodeScholar
2026

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

CVPR 2026

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAg

Cited by 0SourcecodeScholar
2026

Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects

ICRA 2026poster

A deep understanding of kinematic structures is essential for robot motion and interaction with the environment. Such understanding is captured through articulated objects, which are essential for physical simulation, motion planning, and policy learning. However, creating these models, particularly…

2026

PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery

CVPR 2026

Panoramic imagery offers a full 360^\circ field of view and is increasingly common in consumer devices. However, it introduces non-pinhole distortions that challenge joint pose estimation and 3D reconstruction. Existing feed-forward models, built for perspective cameras, generalize poorly to this se

Cited by 0SourcecodeScholar
2026

Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction

ICML 2026poster

Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained generative models as modular priors. However, we identify a critical flaw in prevailing PnP solvers (e.g., based on HQS or Proximal Gradient): they functio…

Cited by 0SourceScholar
2026

SPREAD: Spatial-Physical REasoning via geometry Aware Diffusion

CVPR 2026

Automated 3D scene generation is pivotal for applications spanning virtual reality, digital content creation, and Embodied AI. While computer graphics prioritizes aesthetic layouts, vision and robotics demand scenes that mirror real-world complexity which current data-driven methods struggle to achi

Cited by 0SourceScholar
2026

Towards Sub-second Biological Foundation Model Infrastructure: A Quantized Consistency Diffusion Framework for Molecular Docking

ICML 2026oral

The emergence of Vibe Researching is transforming scientific research into an interactive workflow, where agents orchestrate complex tasks via the Model Context Protocol (MCP). In this ecosystem, scientific tools must evolve from offline simulators into responsive Agent Skills. However, diffusion-ba…

Cited by 0SourceScholar
2026

Unsupervised Multi-Parameter Inverse Solving for Reducing Ring Artifacts in 3D X-Ray CBCT

AAAI 2026technical

Ring artifacts are prevalent in 3D cone-beam computed tomography (CBCT) due to non-ideal responses of X-ray detectors, substantially affecting image quality and diagnostic reliability. Existing state-of-the-art (SOTA) ring artifact reduction (RAR) methods rely on supervised learning with large-scale

Cited by 0SourcePDFScholar
2025

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

CVPR 2025poster

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for diffusion policy. However, their generalization is typically lim…

Cited by 5SourcePDFScholar
2025

BG-Triangle: Bezier Gaussian Triangle for 3D Vectorization and Rendering

CVPR 2025poster

Differentiable rendering enables efficient optimization by allowing gradients to be computed through the rendering process, facilitating 3D reconstruction, inverse rendering and neural scene representation learning. To ensure differentiability, existing solutions approximate or re-formulate traditio…

Cited by 1SourcePDFScholar
2025

Capturing the Unseen: Vision-Free Facial Motion Capture Using Inertial Measurement Units

AAAI 2025technical

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Units (IMUs) as a new sensing modality for facial motion capture. While IMUs have become essential in full-body MoCap for t…

Cited by 0SourcePDFScholar
2025

CryoFastAR: Fast Cryo-EM Ab initio Reconstruction Made Easy

ICCV 2025poster

Pose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in scientific imaging fields like cryo-electron microscopy (cryo-EM) fo…

Cited by 0SourcePDFScholar
2025

DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness

CVPR 2025highlight

A dexterous hand capable of grasping any object is essential for the development of general-purpose embodied intelligent robots. However, due to the high degree of freedom in dexterous hands and the vast diversity of objects, generating high-quality, usable grasping poses in a robust manner is a sig…

2025

DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover

ICCV 2025poster

Handover between a human and a dexterous robotic hand is a fundamental yet challenging task in human-robot collaboration. It requires handling dynamic environments and a wide variety of objects and demands robust and adaptive grasping strategies. However, progress in developing effective dynamic dex…

2025

Discovering Influential Neuron Path in Vision Transformers

ICLR 2025poster

Vision Transformer models exhibit immense power yet remain opaque to human understanding, posing challenges and risks for practical applications. While prior research has attempted to demystify these models through input attribution and neuron role analysis, there's been a notable gap in considerin…

2025

EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment

ICCV 2025poster

Dexterous robotic hands often struggle to generalize effectively in complex environments due to models trained on low-diversity data. However, the real world presents an inherently unbounded range of scenarios. A natural solution is to enable robots learning from experience in complex environments--…

Cited by 0SourcePDFScholar
2025

ExFace: Expressive Facial Control for Humanoid Robots with Diffusion Transformers and Bootstrap Training

IROS 2025

This paper presents a novel Expressive Facial Control (ExFace) method based on Diffusion Transformers, which achieves precise mapping from human facial blendshapes to bionic robot motor control. By incorporating an innovative model bootstrap training strategy, our approach not only generates high-qu

Cited by 2SourceScholar
2025

Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

ICML 2025poster

Diffusion models have emerged as mainstream framework in visual generation. Building upon this success, the integration of Mixture of Experts (MoE) methods has shown promise in enhancing model scalability and performance. In this paper, we introduce Race-DiT, a novel MoE model for diffusion transfor…

Cited by 0SourcePDFScholar
2025

LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing

NeurIPS 2025poster

Lithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality o…

Cited by 0SourcecodeScholar
2025

Moner: Motion Correction in Undersampled Radial MRI with Unsupervised Neural Representation

ICLR 2025spotlight

Motion correction (MoCo) in radial MRI is a particularly challenging problem due to the unpredictability of subject movement. Current state-of-the-art (SOTA) MoCo algorithms often rely on extensive high-quality MR images to pre-train neural networks, which constrains the solution space and leads to…

2025

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

CVPR 2025highlight

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite its promise, real-world datasets often contain noisy labels that can degrade prompt learning performance. In this paper,…

2025

PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding

NeurIPS 2025poster

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress in 3D part understanding, their reliance on untextured geometries and expert-dependent annotation limits scalability and…

Cited by 0SourceScholar
2025

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

AAAI 2025technical

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign language recognition (SLR) and translation (SLT) struggle with dialogue scenes due to limited dataset diversity and the neglec…

2025

SMGDiff: Soccer Motion Generation using Diffusion Probabilistic Models

ICCV 2025poster

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate interactions between the player and the ball. In this paper, we introduce SMGDiff, a novel two-stage framework for generat…

Cited by 0SourcePDFScholar
2025

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

CVPR 2025poster

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance…

Cited by 3SourcePDFScholar
2025

TokMan:Tokenize Manhattan Mask Optimization for Inverse Lithography

NeurIPS 2025poster

Manhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lith…

Cited by 0SourceScholar
2025

Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

ICCV 2025poster

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate progress in kinematic motion generation, they often fail to address the fundamental tension between real-time responsivene…

Cited by 0SourcePDFScholar
2025

TransiT: Transient Transformer for Non-line-of-sight Videography

ICCV 2025poster

High quality and high speed videography using Non-Line-of-Sight (NLOS) imaging benefit autonomous navigation, collision prevention, and post-disaster search and rescue tasks. Current solutions have to balance between the frame rate and image quality. High frame rates, for example, can be achieved by…

Cited by 0SourcePDFScholar
2025

UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control

ICML 2025spotlight

Recent advances in diffusion bridge models leverage Doob’s $h$-transform to establish fixed endpoints between distributions, demonstrating promising results in image translation and restoration tasks. However, these approaches frequently produce blurred or excessively smoothed image details and lack…

2024

A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals

CVPR 2024poster

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from sparse observations to dense full-body motions which endowed inhe…

2024

BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics

CVPR 2024poster

The recently emerging text-to-motion advances have spired numerous attempts for convenient and interactive human motion generation. Yet existing methods are largely limited to generating body motions only without considering the rich two-hand motions let alone handling various conditions like body d…

2024

Content-Aware Radiance Fields: Aligning Model Complexity with Scene Intricacy Through Learned Bitwidth Quantization

ECCV 2024poster

"The recent popular radiance field models, exemplified by Neural Radiance Fields (NeRF), Instant-NGP and 3D Gaussian Splatting, are designed to represent 3D content by that training models for each individual scene. This unique characteristic of scene representation and per-scene training distinguis…

2024

CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy

NeurIPS 2024poster

In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of prot…

Cited by 0SourcePDFScholar
2024

DRACO: A Denoising-Reconstruction Autoencoder for Cryo-EM

NeurIPS 2024poster

Foundation models in computer vision have demonstrated exceptional performance in zero-shot and few-shot tasks by extracting multi-purpose features from large-scale datasets through self-supervised pre-training methods. However, these models often overlook the severe corruption in cryogenic electron…

Cited by 1SourcePDFScholar
2024

Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

NeurIPS 2024poster

Diffusion models have garnered widespread attention in Reinforcement Learning (RL) for their powerful expressiveness and multimodality. It has been verified that utilizing diffusion policies can significantly improve the performance of RL algorithms in continuous control tasks by overcoming the limi…

2024

Guidance with Spherical Gaussian Constraint for Conditional Diffusion

ICML 2024poster

Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step…

2024

HOI-M^3: Capture Multiple Humans and Objects Interaction within Contextual Environment

CVPR 2024highlight

Humans naturally interact with both others and the surrounding multiple objects engaging in various social activities. However recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and objects due to fundamental data scarcity. In this paper we introduc…

Cited by 12SourcePDFScholar
2024

HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting

CVPR 2024poster

We have recently seen tremendous progress in photo-real human modeling and rendering. Yet efficiently rendering realistic human performance and integrating it into the rasterization pipeline remains challenging. In this paper we present HiFi4G an explicit and compact Gaussian-based approach for high…

Cited by 49SourcePDFScholar
2024

I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions

CVPR 2024poster

We are living in a world surrounded by diverse and "smart" devices with rich modalities of sensing ability. Conveniently capturing the interactions between us humans and these objects remains far-reaching. In this paper we present I'm-HOI a monocular scheme to faithfully capture the 3D motions of bo…

Cited by 7SourcePDFScholar
2024

LiveHPS: LiDAR-based Scene-level Human Pose and Shape Estimation in Free Environment

CVPR 2024highlight

For human-centric large-scale scenes fine-grained modeling for 3D human global pose and shape is significant for scene understanding and can benefit many real-world applications. In this paper we present LiveHPS a novel single-LiDAR-based approach for scene-level human pose and shape estimation with…

Cited by 15SourcePDFScholar
2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

NeurIPS 2024poster

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately,…

2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

CVPR 2024poster

We have recently seen tremendous progress in realistic text-to-motion generation. Yet the existing methods often fail or produce implausible motions with unseen text inputs which limits the applications. In this paper we present OMG a novel framework which enables compelling motion generation from z…

2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

IJCAI 2024poster

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and multimodal visual data. Utilizing a teleoperation system, we seamlessly synchronize human-robot hand poses in real time. Th…

2024

VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams

CVPR 2024poster

Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However rendering dynamic long-duration radiance fields on ubiquitous devices remains challenging due to data storage and computational constraints. In this paper we introduce VideoRF the first approach to enable rea…

2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

NeurIPS 2023poster

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a s…

2023

Human-centric Scene Understanding for 3D Large-scale Scenarios

ICCV 2023poster

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset…

Cited by 26PDFcodeScholar
2023

HybridCap: Inertia-Aid Monocular Capture of Challenging Human Motions

AAAI 2023technical

Monocular 3D motion capture (mocap) is beneficial to many applications. The use of a single camera, however, often fails to handle occlusions of different body parts and hence it is limited to capture relatively simple movements. We present a light-weight, hybrid mocap technique called HybridCap tha…

2023

IKOL: Inverse Kinematics Optimization Layer for 3D Human Pose and Shape Estimation via Gauss-Newton Differentiation

AAAI 2023technical

This paper presents an inverse kinematic optimization layer (IKOL) for 3D human pose and shape estimation that leverages the strength of both optimization- and regression-based methods within an end-to-end framework. IKOL involves a nonconvex optimization that establishes an implicit mapping from an…

2023

MotionGPT: Human Motion as a Foreign Language

NeurIPS 2023poster

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multimodal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to human language, often perc…

2023

NeMF: Inverse Volume Rendering with Neural Microflake Field

ICCV 2023poster

Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering.Recent approaches adopt the emerging implicit scene representations and have shown impressive results.However, they unanimous…

Cited by 26PDFcodeScholar
2023

NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions

ICCV 2023oral

We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The…

Cited by 82PDFcodeScholar
2023

Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos

CVPR 2023poster

The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has inspired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating free-view videos (FVVs) are restricted to either offline rendering or are capable…

Cited by 66SourcePDFScholar
2023

NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions

CVPR 2023poster

Humans constantly interact with objects in daily life tasks. Capturing such processes and subsequently conducting visual inferences from a fixed viewpoint suffers from occlusions, shape and texture ambiguities, motions, etc. To mitigate the problem, it is essential to build a training dataset that c…

2023

Relightable Neural Human Assets From Multi-View Gradient Illuminations

CVPR 2023poster

Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most existing human datasets only provide multi-view human images captured under the same illumination. Although valuable for mode…

2023

ScalableMap: Scalable Map Learning for Online Long-Range Vectorized HD Map Construction

CoRL 2023poster

We propose a novel end-to-end pipeline for online long-range vectorized high-definition (HD) map construction using on-board camera sensors. The vectorized representation of HD maps, employing polylines and polygons to represent map elements, is widely used by downstream tasks. However, previous sch…

Cited by 18SourcecodeScholar
2023

StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset

IJCAI 2023poster

Modeling and capturing the 3D spatial arrangement of the human and the object is the key to perceiving 3D human-object interaction from monocular images. In this work, we propose to use the Human-Object Offset between anchors which are densely sampled from the surface of human mesh and object mesh t…

2023

Unsupervised Polychromatic Neural Representation for CT Metal Artifact Reduction

NeurIPS 2023poster

Emerging neural reconstruction techniques based on tomography (e.g., NeRF, NeAT, and NeRP) have started showing unique capabilities in medical imaging. In this work, we present a novel Polychromatic neural representation (Polyner) to tackle the challenging problem of CT imaging when metallic implant…

2023

Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDAR

AAAI 2023technical

Depth estimation is usually ill-posed and ambiguous for monocular camera-based 3D multi-person pose estimation. Since LiDAR can capture accurate depth information in long-range scenes, it can benefit both the global localization of individuals and the 3D pose estimation by providing rich geometry fe…

2022

Anisotropic Fourier Features for Neural Image-Based Rendering and Relighting

AAAI 2022technical

Recent neural rendering techniques have greatly benefited image-based modeling and relighting tasks. They provide a continuous, compact, and parallelable representation by modeling the plenoptic function as multilayer perceptrons (MLPs). However, vanilla MLPs suffer from spectral biases on multidime…

Cited by 7SourcePDFScholar
2022

Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-Time

CVPR 2022oral

Implicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can be achieved with smart data structures, e.g., PlenOctree. In this paper, we present a novel Fourier PlenOctree (FPO) te…

Cited by 178PDFScholar
2022

HSC4D: Human-Centered 4D Scene Capture in Large-Scale Indoor-Outdoor Space Using Wearable IMUs and LiDAR

CVPR 2022poster

We propose Human-centered 4D Scene Capture (HSC4D) to accurately and efficiently create a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, and rich interactions between humans and environments. Using only body-mounted IMUs and LiDAR, HSC4D is space-free wit…

Cited by 36PDFcodeScholar
2022

HumanNeRF: Efficiently Generated Human Radiance Field From Sparse Inputs

CVPR 2022poster

Recent neural human representations can produce high-quality multi-view rendering but require using dense multi-view inputs and costly training. They are hence largely limited to static models as training each frame is infeasible. We present HumanNeRF - a neural representation with efficient general…

Cited by 227PDFScholar
2022

LiDARCap: Long-Range Marker-Less 3D Human Motion Capture With LiDAR Point Clouds

CVPR 2022poster

Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captured by LiDAR at a much longer range to overcome this limitation. Our dataset also includes the ground truth human motions…

Cited by 62PDFScholar
2022

NeuralHOFusion: Neural Volumetric Rendering Under Human-Object Interactions

CVPR 2022poster

4D modeling of human-object interactions is critical for numerous applications. However, efficient volumetric capture and rendering of complex interaction scenarios, especially from sparse inputs, remain challenging. In this paper, we propose NeuralHOFusion, a neural approach for volumetric human-ob…

Cited by 50PDFScholar
2021

Few-shot Neural Human Performance Rendering from Sparse RGBD Videos

IJCAI 2021poster

Recent neural rendering approaches for human activities achieve remarkable view synthesis results, but still rely on dense input views or dense training with all the capture frames, leading to deployment difficulty and inefficient training overload. However, existing advances will be ill-posed if th…

Cited by 17SourcePDFScholar
2021

GNeRF: GAN-Based Neural Radiance Field Without Posed Camera

ICCV 2021poster

We introduce GNeRF, a framework to marry Generative Adversarial Networks (GAN) with Neural Radiance Field (NeRF) reconstruction for the complex scenarios with unknown and even randomly initialized camera poses. Recent NeRF-based advances have gained popularity for remarkable realistic novel view syn…

Cited by 223PDFcodeScholar
2021

MVSNeRF: Fast Generalizable Radiance Field Reconstruction From Multi-View Stereo

ICCV 2021poster

We present MVSNeRF, a novel neural rendering approach that can efficiently reconstruct neural radiance fields for view synthesis. Unlike prior works on neural radiance fields that consider per-scene optimization on densely captured images, we propose a generic deep neural network that can reconstruc…

Cited by 907PDFcodeScholar
2021

Neural Video Portrait Relighting in Real-Time via Consistency Modeling

ICCV 2021poster

Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result under dynamic illuminations from monocular RGB stream, suffering from the lack of video consistency supervision. In this p…

Cited by 47PDFcodeScholar
2021

PIANO: A Parametric Hand Bone Model from Magnetic Resonance Imaging

IJCAI 2021poster

Hand modeling is critical for immersive VR/AR, action understanding, or human healthcare. Existing parametric models account only for hand shape, pose, or texture, without modeling the anatomical attributes like bone, which is essential for realistic hand biomechanics analysis. In this paper, we pre…

2020

A Neural Rendering Framework for Free-Viewpoint Relighting

CVPR 2020poster

We present a novel Relightable Neural Renderer (RNR) for simultaneous view synthesis and relighting using multi-view image inputs. Existing neural rendering (NR) does not explicitly model the physical rendering process and hence has limited capabilities on relighting. RNR instead models image format…

Cited by 60PDFcodeScholar
2020

Geometric Structure Based and Regularized Depth Estimation From 360 Indoor Imagery

CVPR 2020poster

Motivated by the correlation between the depth and the geometric structure of a 360 indoor image, we propose a novel learning-based depth estimation framework that leverages the geometric structure of a scene to conduct depth estimation. Specifically, we represent the geometric structure of an indoo…

Cited by 86PDFScholar
2020

Spatial-Angular Interaction for Light Field Image Super-Resolution

ECCV 2020poster

Light field (LF) cameras record both intensity and directions of light rays, and capture scenes from a number of viewpoints. Both information within each perspective (i.e., spatial information) and among different perspectives (i.e., angular information) is beneficial to image super-resolution (SR).…

2019

Photo-Realistic Facial Details Synthesis From Single Image

ICCV 2019oral

We present a single-image 3D face synthesis technique that can handle challenging facial expressions while recovering fine geometric details. Our technique employs expression analysis for proxy face geometry generation and combines supervised and unsupervised learning for facial detail synthesis. On…

Cited by 131PDFcodeScholar
2018

Automatic 3D Indoor Scene Modeling From Single Panorama

CVPR 2018poster

We describe a system that automatically extracts 3D geometry of an indoor scene from a single 2D panorama. Our system recovers the spatial layout by finding the floor, walls, and ceiling; it also recovers shapes of typical indoor objects such as furniture. Using sampled perspective sub-views, we ext…

Cited by 67SourcePDFScholar
2018

Gaze Prediction in Dynamic 360° Immersive Videos

CVPR 2018poster

This paper explores gaze prediction in dynamic $360^circ$ immersive videos, emph{i.e.}, based on the history scan path and VR contents, we predict where a viewer will look at an upcoming time. To tackle this problem, we first present the large-scale eye-tracking in dynamic VR scene dataset. Our data…

2018

Learning to Dodge A Bullet: Concyclic View Morphing via Deep Learning

ECCV 2018poster

The bullet-time effect, presented in feature film ``The Matrix", has been widely adopted in feature films and TV commercials to create an amazing stopping-time illusion. Producing such visual effects, however, typically requires using a large number of cameras/images surrounding the subject. In this…

Cited by 13SourcePDFScholar
2018

Sparse Photometric 3D Face Reconstruction Guided by Morphable Models

CVPR 2018poster

We present a novel 3D face reconstruction technique that leverages sparse photometric stereo (PS) and latest advances on face registration / modeling from a single image. We observe that 3D morphable faces approach provides a reasonable geometry proxy for light position calibration. Specifically, we…

Cited by 41SourcePDFScholar
2017

A new calibration technique for multi-camera systems of limited overlapping field-of-views

IROS 2017poster

State-of-the-art calibration methods typically choose to use a checkerboard as the calibration target for its simplicity and robustness. They however require the complete checkerboard be captured to break symmetry. More recent multi-camera systems such as Google Jump, Jaunt, and camera arrays have l…

Cited by 27SourceScholar
2015

Ambient Occlusion via Compressive Visibility Estimation

CVPR 2015poster

There has been emerging interest on recovering traditionally challenging intrinsic scene properties. In this paper, we present a novel computational imaging solution for recovering the ambient occlusion (AO) map of an object. AO measures how much light from all different directions can reach a surfa…

Cited by 10SourcePDFScholar