← Search

YAO YAO

54 accepted papers

2026

Anime-Ready: Controllable 3D Anime Character Generation with Body-Aligned Component-Wise Garment Modeling

ICLR 2026poster

3D anime character generation has become increasingly important in digital entertainment, including animation production, virtual reality, gaming, and virtual influencers. Unlike realistic human modeling, anime-style characters require exaggerated proportions, stylized surface details, and artistica…

Cited by 0SourceScholar
2026

ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes

ICLR 2026poster

Gaussian Splatting (GS) enables immersive rendering, but realistic 3D object–scene composition remains challenging. Baked appearance and shadow information in GS radiance fields cause inconsistencies when combining objects and scenes. Addressing this requires relightable object reconstruction and sc…

Cited by 0SourcecodeScholar
2026

LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging

CVPR 2026

3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To address this, we propose LiteVGGT

Cited by 0SourcecodeScholar
2026

Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance

CVPR 2026

We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requires only a pressure mat, eliminating the need for specialized lighting setups, cameras, or wearable devices, making it

Cited by 0SourcecodeScholar
2026

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

CVPR 2026

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of large-scale, high-quality training data. While several datasets pr

Cited by 0SourcecodeScholar
2026

TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond

CVPR 2026

Prevailing 3D texture generation methods, which often rely on multi-view fusion, are frequently hindered by inter-view inconsistencies and incomplete coverage of complex surfaces, limiting the fidelity and completeness of the generated content. To overcome these challenges, we introduce TEXTRIX, a n

Cited by 0SourcecodeScholar
2025

4D Diffusion for Dynamic Protein Structure Prediction with Reference and Motion Guidance

AAAI 2025technical

Protein structure prediction is pivotal for understanding the structure-function relationship of proteins, advancing biological research, and facilitating pharmaceutical development and experimental design. While deep learning methods and the expanded availability of experimental 3D protein structur…

Cited by 0SourcePDFScholar
2025

An End-to-End Framework for Modeling Pneumatic Soft Robots Based on Differentiable Finite Element Methods

RA-L 2025

Soft robots present significant modelling challenges due to their non-linearity, complex dynamics and potentially intricate geometries. These difficulties in accurate system identification and dynamics modelling limit their applications in precise robotics tasks. Prior modelling approaches typically

Cited by 0SourceScholar
2025

Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

ACL 2025long

This paper investigates the faithfulness of multimodal large language model (MLLM) agents in a graphical user interface (GUI) environment, aiming to address the research question of whether multimodal GUI agents can be distracted by environmental context. A general scenario is proposed where both th…

2025

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

NeurIPS 2025poster

Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dra…

Cited by 0SourceScholar
2025

FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

CVPR 2025poster

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstructi…

2025

Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priors

ICLR 2025poster

3D Gaussian Splatting (3DGS) has achieved excellent rendering quality with fast training and rendering speed. However, its optimization process lacks explicit geometric constraints, leading to suboptimal geometric reconstruction in regions with sparse or no observational input views. In this work, w…

Cited by 0SourcePDFScholar
2025

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

ICLR 2025poster

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend…

2025

JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation

AAAI 2025technical

With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving precise control over multi-track generation remains an open challenge. While existing models excel in directly generating mul…

Cited by 12SourcePDFScholar
2025

JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning

AAAI 2025technical

Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts. However, input text prompts may not precisely capture user requirements, particularly when the objective is to generate…

Cited by 2SourcePDFScholar
2025

Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh

CVPR 2025poster

Neural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and i…

Cited by 10SourcePDFScholar
2025

Matrix3D: Large Photogrammetry Model All-in-One

CVPR 2025highlight

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, suc…

2025

XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks. However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in r

2024

Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

ECCV 2024poster

"In this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in current human generative techniques. The methodology utilizes the SMPL(Skinned Multi-Person Linear) mod…

2024

Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video

ICLR 2024poster

In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration…

2024

Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion

CVPR 2024poster

Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS) or a direct 3D diffusion model trained on limited 3D data losing genera…

Cited by 33SourcePDFScholar
2024

Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

NeurIPS 2024poster

Generating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images,…

Cited by 35SourcePDFScholar
2024

EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head

ECCV 2024poster

"We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address…

2024

GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment

ACL 2024findings

The burgeoning size of Large Language Models (LLMs) has led to enhanced capabilities in generating responses, albeit at the expense of increased inference times and elevated resource demands. Existing methods of acceleration, predominantly hinged on knowledge distillation, generally necessitate fine…

2024

Gaussian-Flow: 4D Reconstruction with Dynamic 3D Gaussian Particle

CVPR 2024highlight

We introduce Gaussian-Flow a novel point-based approach for fast dynamic scene reconstruction and real-time rendering from both multi-view and monocular videos. In contrast to the prevalent NeRF-based approaches hampered by slow training and rendering speeds our approach harnesses recent advancement…

Cited by 99SourcePDFScholar
2024

GaussianPro: 3D Gaussian Splatting with Progressive Propagation

ICML 2024poster

3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain tex…

2024

Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360°

ECCV 2024poster

"Creating a 360◦ parametric model of a human head is a very challenging task. While recent advancements have demonstrated the efficacy of leveraging synthetic data for building such parametric head models, their performance remains inadequate in crucial areas such as expression-driven animation, hai…

2024

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

ICLR 2024poster

We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense moda…

Cited by 10SourcePDFScholar
2024

Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have rapidly advanced and demonstrated impressive capabilities. In-Context Learning (ICL) and Parameter-Efficient Fine-Tuning (PEFT) are currently two mainstream methods for augmenting LLMs to downstream tasks. ICL typically constructs a few-shot learning scenario, eithe…

2024

Rotated Orthographic Projection for Self-Supervised 3D Human Pose Estimation

ECCV 2024poster

"Reprojection consistency is widely used for self-supervised 3D human pose estimation. However, few efforts have been made to address the inherent limitations of reprojection consistency. Lacking camera parameters and absolute position, self-supervised methods map 3D poses to 2D using orthographic p…

Cited by 0SourcePDFScholar
2024

Stereo Risk: A Continuous Modeling Approach to Stereo Matching

ICML 2024oral

We introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing…

Cited by 8SourcePDFScholar
2023

Coarse2Fine: Local Consistency Aware Re-prediction for Weakly Supervised Object Localization

AAAI 2023technical

Weakly supervised object localization aims to localize objects of interest by using only image-level labels. Existing methods generally segment activation map by threshold to obtain mask and generate bounding box. However, the activation map is locally inconsistent, i.e., similar neighboring pixels…

Cited by 10SourcePDFScholar
2023

NeILF++: Inter-Reflectable Light Fields for Geometry and Material Estimation

ICCV 2023poster

We present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one ne…

Cited by 56PDFScholar
2023

Stochastic Methods for AUC Optimization subject to AUC-based Fairness Constraints

AISTATS 2023poster

As machine learning being used increasingly in making high-stakes decisions, an arising challenge is to avoid unfair AI systems that lead to discriminatory decisions for protected population. A direct approach for obtaining a fair predictive model is to train the model through optimizing its predict…

Cited by 7SourcePDFScholar
2022

Critical Regularizations for Neural Surface Reconstruction in the Wild

CVPR 2022poster

Neural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that…

Cited by 54PDFScholar
2022

NeILF: Neural Incident Light Field for Physically-Based Material Estimation

ECCV 2022poster

"We present a differentiable rendering framework for material and lighting estimation from multi-view images and a reconstructed geometry. In the framework, we represent scene lightings as the Neural Incident Light Field (NeILF) and material properties as the surface BRDF modelled by multi-layer per…

Cited by 111SourcePDFScholar
2021

Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation

ICRA 2021poster

Model-based deep reinforcement learning has achieved success in various domains that require high sample efficiencies, such as Go and robotics. However, there are some remaining issues, such as planning efficient explorations to learn more accurate dynamic models, evaluating the uncertainty of the l…

Cited by 29SourcecodeScholar
2020

ASLFeat: Learning Local Features of Accurate Shape and Localization

CVPR 2020poster

This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to ac…

Cited by 379PDFcodeScholar
2020

BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo Networks

CVPR 2020poster

While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensiv…

Cited by 534PDFcodeScholar
2020

KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering

CVPR 2020oral

Temporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the…

Cited by 99PDFcodeScholar
2019

ContextDesc: Local Descriptor Augmentation With Cross-Modality Context

CVPR 2019oral

Most existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment…

Cited by 315PDFcodeScholar
2019

Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface

CVPR 2019poster

We present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the paramete…

Cited by 21PDFScholar
2019

Non-local Self-attention Structure for Function Approximation in Deep Reinforcement Learning

ICASSP 2019accepted

Reinforcement learning is a framework to make sequential decisions. The combination with deep neural networks further improves the ability of this framework. Convolutional nerual networks make it possible to make sequential decisions based on raw pixels information directly and make reinforcement le…

Cited by 0SourceScholar
2019

Recurrent MVSNet for High-Resolution Multi-View Stereo Depth Inference

CVPR 2019poster

Deep learning has recently demonstrated its excellent performance for multi-view stereo (MVS). However, one major limitation of current learned MVS approaches is the scalability: the memory-consuming cost volume regularization makes the learned MVS hard to be applied to high-resolution scenes. In th…

Cited by 706PDFcodeScholar
2018

GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints

ECCV 2018poster

Learned local descriptors based on Convolutional Neural Networks (CNNs) have achieved significant improvements on patch-based benchmarks, whereas not having demonstrated strong generalization ability on recent benchmarks of image-based 3D reconstruction. In this paper, we mitigate this limitation by…

Cited by 216SourcePDFScholar
2018

MVSNet: Depth Inference for Unstructured Multi-view Stereo

ECCV 2018poster

We present an end-to-end deep learning architecture for depth map inference from multi-view images. In the network, we first extract deep visual image features, and then build the 3D cost volume upon the reference camera frustum via the differentiable homography warping. Next, we apply 3D convolutio…

2018

Reconstructing Thin Structures of Manifold Surfaces by Integrating Spatial Curves

CVPR 2018poster

The manifold surface reconstruction in multi-view stereo often fails in retaining thin structures due to incomplete and noisy reconstructed point clouds. In this paper, we address this problem by leveraging spatial curves. The curve representation in nature is advantageous in modeling thin and elong…

Cited by 39SourcePDFScholar