← Search

Hao Zhu

93 accepted papers

2026

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

ICLR 2026poster

Agents are predominantly evaluated and optimized via task success metrics, which are coarse, rely on manual design from experts, and fail to reward intermediate emergent behaviors. We propose AutoLibra, a framework for agent evaluation, that transforms open-ended human feedback e.g. “If you find tha…

Cited by 0SourcecodeScholar
2026

ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes

ICLR 2026poster

Gaussian Splatting (GS) enables immersive rendering, but realistic 3D object–scene composition remains challenging. Baked appearance and shadow information in GS radiance fields cause inconsistencies when combining objects and scenes. Addressing this requires relightable object reconstruction and sc…

Cited by 0SourcecodeScholar
2026

CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Image

CVPR 2026

Single-view 3D human reconstruction has garnered significant attention in recent years. Despite numerous advancements, prior research has concentrated on reconstructing 3D models from clear, close-up images of individual subjects, often yielding subpar results in the more prevalent multi-person scen

Cited by 0SourceScholar
2026

DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations

CVPR 2026

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these diffusion models realizes high-fidelity disentangled control betw

Cited by 0SourceScholar
2026

Hierarchical Action Learning for Weakly-Supervised Action Segmentation

CVPR 2026

Humans perceive actions through key transitions that structure actions across multiple abstraction levels, whereas machines, relying on visual features, tend to over-segment. This highlights the difficulty of enabling hierarchical reasoning in video understanding. Interestingly, we observe that lowe

Cited by 0SourcecodeScholar
2026

Hierarchical Direction Perception via Atomic Dot-Product Operators for Rotation-Invariant Point Clouds Learning

AAAI 2026technical

Point cloud processing has become a cornerstone technology in many 3D vision tasks. However, arbitrary rotations introduce variations in point cloud orientations, posing a long-standing challenge for effective representation learning. The core of this issue is the disruption of the point cloud

Cited by 0SourcePDFScholar
2026

Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance

CVPR 2026

We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requires only a pressure mat, eliminating the need for specialized lighting setups, cameras, or wearable devices, making it

Cited by 0SourcecodeScholar
2026

RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection

AAAI 2026technical

Unsupervised industrial anomaly detection requires accurately identifying defects without labeled data. Traditional autoencoder-based methods often struggle with incomplete anomaly suppression and loss of fine details, as their single-pass decoding fails to effectively handle anomalies with varying

Cited by 0SourcePDFScholar
2026

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

CVPR 2026

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of large-scale, high-quality training data. While several datasets pr

Cited by 0SourcecodeScholar
2026

Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature Space

AAAI 2026technical

Implicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of INR—defined by the range of functions the neural network ca

Cited by 0SourcePDFScholar
2026

TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond

CVPR 2026

Prevailing 3D texture generation methods, which often rely on multi-view fusion, are frequently hindered by inter-view inconsistencies and incomplete coverage of complex surfaces, limiting the fidelity and completeness of the generated content. To overcome these challenges, we introduce TEXTRIX, a n

Cited by 0SourcecodeScholar
2026

Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning

ICLR 2026poster

The composition of specialized tools offers a powerful approach for complex visual reasoning, particularly for tasks involving 3D spatial understanding. However, existing visual programming methods are often constrained by fixed toolsets or offline tool induction, which leads to suboptimal solutions…

Cited by 0SourcecodeScholar
2026

UIKA: Fast Universal Head Avatar from Pose-Free Images

CVPR 2026

We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view captures, and smartphone-captured videos. Unlike the traditional avatar method, which requires a studio-level multi-view capture system and reconstructs a

Cited by 0SourcecodeScholar
2026

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

ICML 2026poster

Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-…

Cited by 0SourceScholar
2025

BiLoRA: Almost-Orthogonal Parameter Spaces for Continual Learning

CVPR 2025poster

Continual learning requires models to learn tasks sequentially while maintaining a delicate balance between stability (retaining knowledge of previous tasks) and plasticity (adapting to new tasks). A key challenge is preventing interference between tasks - where learning new tasks degrades performan…

2025

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time

ICML 2025poster

Aligning large language models with humans is challenging due to the inherently multifaceted nature of preference feedback. While existing approaches typically frame this as a multi-objective optimization problem, they often overlook how humans actually make decisions. Research on bounded rationalit…

Cited by 0SourcePDFScholar
2025

CrossSpectra: Exploiting Cross-Layer Smoothness for Parameter-Efficient Fine-Tuning

NeurIPS 2025poster

Parameter-efficient fine-tuning (PEFT) is essential for adapting large foundation models without excessive storage cost. However, current approaches such as LoRA treat each layer’s adaptation independently, overlooking correlations across layers. This independence causes the number of trainable para…

Cited by 0SourceScholar
2025

Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstruction

CVPR 2025poster

Recent advancements in generalizable novel view synthesis have achieved impressive quality through interpolation between nearby views. However, rendering high-resolution images remains computationally intensive due to the need for dense sampling of all rays. Recognizing that natural scenes are typic…

2025

EgoNormia: Benchmarking Physical-Social Norm Understanding

ACL 2025finding

Human activity is moderated by norms; however, supervision for normative reasoning is sparse, particularly where norms are physically- or socially-grounded. We thus present EgoNormia \lVert 𝜖 \rVert, comprising 1,853 (200 for EgoNormia-verified) multiple choice questions (MCQs) grounded within ego-c…

2025

Exact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic Segmentation

CVPR 2025highlight

Automated crop mapping through Satellite Image Time Series (SITS) has emerged as a crucial avenue for agricultural monitoring and management. However, due to the low resolution and unclear parcel boundaries, annotating pixel-level masks is exceptionally complex and time-consuming in SITS. This paper…

2025

FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

CVPR 2025poster

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstructi…

2025

From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach

AAAI 2025technical

In this paper, we present CAD2Program, a new method for reconstructing 3D parametric models from 2D CAD drawings. Our proposed method is inspired by recent successes in vision-language models (VLMs), and departs from traditional methods which rely on task-specific data representations and/or algorit…

2025

GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis

CVPR 2025poster

Neural Radiance Fields (NeRF) have transformed novel view synthesis by modeling scene-specific volumetric representations directly from images. While generalizable NeRF models can generate novel views across unknown scenes by learning latent ray representations, their performance heavily depends on…

2025

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

ICLR 2025poster

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend…

2025

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

CVPR 2025poster

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the…

2025

Improving Zero-Shot Adversarial Robustness in Vision-Language Models by Closed-form Alignment of Adversarial Path Simplices

ICML 2025spotlight

Vision-Language Models (VLMs) such as CLIP excel at zero-shot classification due to large-scale pre-training but are vulnerable to adversarial examples. Adversarial fine-tuning robustifies zero-shot models by aligning prediction scores of individual adversaries with their clean counterparts, which t…

Cited by 0SourcePDFScholar
2025

LEVIS: Large Exact Verifiable Input Spaces for Neural Networks

ICML 2025poster

The robustness of neural networks is crucial in safety-critical applications, where identifying a reliable input space is essential for effective model selection, robustness evaluation, and the development of reliable control strategies. Most existing robustness verification methods assess the worst…

Cited by 0SourcePDFScholar
2025

Machine Unlearning via Task Simplex Arithmetic

NeurIPS 2025poster

As foundation Vision-Language Models (VLMs) unlock fine-tuning on smaller datasets while leveraging large-scale pre-training data, machine unlearning becomes critical in addressing privacy concerns and regulatory compliance. Task vector, representing the difference between parameters of models fine-…

Cited by 0SourceScholar
2025

Mind the Gap: Static and Interactive Evaluations of Large Audio Models

ACL 2025long

As AI chatbots become ubiquitous, voice interaction presents a compelling way to enable rapid, high-bandwidth communication for both semantic and social signals. This has driven research into Large Audio Models (LAMs) to power voice-native experiences. However, aligning LAM development with user goa…

2025

Mitigating Ambiguities in 3D Classification with Gaussian Splatting

CVPR 2025poster

3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective obje…

Cited by 0SourcePDFScholar
2025

Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity Modeling

NeurIPS 2025poster

Periodicity, as one of the most important basic characteristics, lays the foundation for facilitating structured knowledge acquisition and systematic cognitive processes within human learning paradigms. However, the potential flaws of periodicity modeling in Transformer affect the learning efficienc…

Cited by 0SourceScholar
2025

Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models

NeurIPS 2025poster

Learning from self-sampled data and sparse environmental feedback remains a fundamental challenge in training self-evolving agents. Temporal credit assignment mitigates this issue by transforming sparse feedback into dense supervision signals. However, previous approaches typically depend on domain-…

Cited by 0SourceScholar
2025

Robustifying Zero-Shot Vision Language Models by Subspaces Alignment

ICCV 2025poster

Vision-Language Models (VLMs) enjoy strong zero-shot performance but are vulnerable to adversarial attacks posing security risks. Adversarially robust fine-tuning enhances zero-shot robustness on new datasets while preserving the natural performance of pre-trained VLMs. However, prior methods use sa…

Cited by 0SourcePDFScholar
2025

SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning

NeurIPS 2025spotlight

How to design reinforcement learning (RL) tasks that effectively unleash the reasoning capability of large language models (LLMs) remains an open question. Existing RL tasks (e.g., math, programming, and constructing reasoning tasks) suffer from three key limitations: (1) Scalability. They rely heav…

Cited by 0SourceScholar
2025

SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation

NAACL 2025system demonstrations

Social simulation through large language model (LLM) agents is a promising approach to explore and validate social science hypotheses.We present SOTOPIA-S4, a fast, flexible, and scalable social simulation system that addresses the technical barriers of current frameworks while enabling practitioner…

2025

SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions

NeurIPS 2025poster

Humans continuously infer the states, goals, and behaviors of others by perceiving their surroundings in dynamic, real-world social interactions. However, most Theory of Mind (ToM) benchmarks only evaluate static, text-based scenarios, which have a significant gap compared to real interactions. We p…

Cited by 0SourceScholar
2025

SpatialLM: Training Large Language Models for Structured Indoor Modeling

NeurIPS 2025poster

SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit…

Cited by 0SourceScholar
2025

TeRA: Rethinking Text-guided Realistic 3D Avatar Generation

ICCV 2025poster

Efficient 3D avatar creation is a significant demand in the metaverse, film/game, AR/VR, etc. In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach…

Cited by 0SourcePDFScholar
2025

pFedMxF: Personalized Federated Class-Incremental Learning with Mixture of Frequency Aggregation

CVPR 2025poster

Federated learning (FL) has emerged as a promising paradigm for privacy-preserving collaborative machine learning. However, extending FL to class incremental learning settings introduces three key challenges: 1) spatial heterogeneity due to non-IID data distributions across clients, 2) temporal hete…

Cited by 0SourcePDFScholar
2024

A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields

AAAI 2024technical

Recent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light mo…

Cited by 2SourcePDFScholar
2024

Batch Normalization Alleviates the Spectral Bias in Coordinate Networks

CVPR 2024poster

Representing signals using coordinate networks dominates the area of inverse problems recently and is widely applied in various scientific computing tasks. Still there exists an issue of spectral bias in coordinate networks limiting the capacity to learn high-frequency components. This problem is ca…

Cited by 9SourcePDFScholar
2024

DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

ACL 2024findings

How to evaluate the coding abilities of Large Language Models (LLMs) remains an open question. We find that existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of LLMs.To address the knowledge gap, we propose a new benchmark…

2024

Exploiting Depth Priors for Few-Shot Neural Radiance Field Reconstruction

RA-L 2024

The performance of neural radiance field technologies deteriorates rapidly when sparse views are used as input. In this paper, we propose a simulated viewpoint enhancement for surface reconstruction that extracts diverse geometric features from the depth to address this limitation. We design a novel

Cited by 0SourceScholar
2024

FINER: Flexible Spectral-bias Tuning in Implicit NEural Representation by Variable-periodic Activation Functions

CVPR 2024poster

Implicit Neural Representation (INR) which utilizes a neural network to map coordinate inputs to corresponding attributes is causing a revolution in the field of signal processing. However current INR techniques suffer from a restricted capability to tune their supported frequency set resulting in i…

Cited by 34SourcePDFScholar
2024

MISA: MIning Saliency-Aware Semantic Prior for Box Supervised Instance Segmentation

IJCAI 2024poster

Box supervised instance segmentation (BSIS) aims to achieve an effective trade-off between annotation costs and model performance by solely relying on bounding box annotations during training process. However, we observe that BSIS model is bottlenecked by the intricate objective under limited guidan…

Cited by 2SourcePDFScholar
2024

Neural Poisson Solver: A Universal and Continuous Framework for Natural Signal Blending

ECCV 2024poster

"Implicit Neural Representation (INR) has become a popular method for representing visual signals (, 2D images and 3D scenes), demonstrating promising results in various downstream applications. Given its potential as a medium for visual signals, exploring the development of a neural blending method…

Cited by 0SourcePDFScholar
2024

PaintHuman: Towards High-Fidelity Text-to-3D Human Texturing via Denoised Score Distillation

AAAI 2024technical

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (e.g., SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide inaccurate gradient directions under the weak diffusion g…

2024

Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing

ECCV 2024poster

"In this paper, we present a novel differentiable point-based rendering framework to achieve photo-realistic relighting. To make the reconstructed scene relightable, we enhance vanilla 3D Gaussians by associating extra properties, including normal vectors, BRDF parameters, and incident lighting from…

Cited by 133SourcePDFScholar
2024

SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents

ACL 2024long

Humans learn social skills through both imitation and social interaction. This social learning process is largely understudied by existing research on building language agents. Motivated by this gap, we propose an interactive learning method, SOTOPIA-π, that improves the social intelligence of langu…

2024

SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

ICLR 2024spotlight

*Humans are social beings*; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and…

Cited by 148SourcePDFScholar
2024

STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians

ECCV 2024poster

"Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In this work, we propose STAG4D, a novel framework that combines pre-trained diffusio…

Cited by 46SourcePDFScholar
2024

WebArena: A Realistic Web Environment for Building Autonomous Agents

ICLR 2024poster

With advances in generative AI, there is now potential for autonomous agents to manage daily tasks via natural language commands. However, current agents are primarily created and tested in simplified synthetic environments, leading to a disconnect with real-world scenarios. In this paper, we build…

2023

COBRA Frames: Contextual Reasoning about Effects and Harms of Offensive Statements

ACL 2023findings

Warning: This paper contains content that may be offensive or upsetting. Understanding the harms and offensiveness of statements requires reasoning about the social and situational context in which statements are made. For example, the utterance “your English is very good” may implicitly signal an i…

2023

CelebV-Text: A Large-Scale Facial Text-Video Dataset

CVPR 2023poster

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, a…

2023

Computational Language Acquisition with Theory of Mind

ICLR 2023poster

Unlike current state-of-the-art language models, young children actively acquire language through interactions with their surrounding environment and caretakers. One mechanism that has been argued to be critical to language learning is the ability to infer the mental states of other agents in social…

2023

DINER: Disorder-Invariant Implicit Neural Representation

CVPR 2023highlight

Implicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that suc…

2023

EXCALIBUR: Encouraging and Evaluating Embodied Exploration

CVPR 2023poster

Experience precedes understanding. Humans constantly explore and learn about their environment out of curiosity, gather information, and update their models of the world. On the other hand, machines are either trained to learn passively from static and fixed datasets, or taught to complete specific…

Cited by 17SourcePDFScholar
2023

Hierarchical Prompting Assists Large Language Model on Web Navigation

EMNLP 2023short findings

Large language models (LLMs) struggle on processing complicated observations in interactive decision making. To alleviate this issue, we propose a simple hierarchical prompting approach. Diverging from previous prompting approaches that always put the full observation (a web page) to the prompt, we…

Cited by 0SourcecodeScholar
2023

High-Fidelity 3D Face Generation From Natural Language Descriptions

CVPR 2023poster

Synthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue the major obstacle lies in 1) the lack of high-quality 3D fa…

2023

Mitigating the Popularity Bias of Graph Collaborative Filtering: A Dimensional Collapse Perspective

NeurIPS 2023spotlight

Graph-based Collaborative Filtering (GCF) is widely used in personalized recommendation systems. However, GCF suffers from a fundamental problem where features tend to occupy the embedding space inefficiently (by spanning only a low-dimensional subspace). Such an effect is characterized in GCF by th…

Cited by 27SourcePDFScholar
2023

RAFaRe: Learning Robust and Accurate Non-parametric 3D Face Reconstruction from Pseudo 2D&3D Pairs

AAAI 2023technical

We propose a robust and accurate non-parametric method for single-view 3D face reconstruction (SVFR). While tremendous efforts have been devoted to parametric SVFR, a visible gap still lies between the result 3D shape and the ground truth. We believe there are two major obstacles: 1) the representat…

2023

Spectral Feature Augmentation for Graph Contrastive Learning and Beyond

AAAI 2023technical

Although augmentations (e.g., perturbation of graph edges, image crops) boost the efficiency of Contrastive Learning (CL), feature level augmentation is another plausible, complementary yet not well researched strategy. Thus, we present a novel spectral feature argumentation for contrastive learni…

2023

Transductive Few-Shot Learning With Prototype-Based Label Propagation by Iterative Graph Refinement

CVPR 2023poster

Few-shot learning (FSL) is popular due to its ability to adapt to novel classes. Compared with inductive few-shot learning, transductive models typically perform better as they leverage all samples of the query set. The two existing classes of methods, prototype-based and graph-based, have the disad…

2022

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

ECCV 2022poster

"Large-scale datasets played an indispensable role in the recent success of face generation/editing and significantly facilitate the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for face-…

2022

Detailed Facial Geometry Recovery from Multi-View Images by Learning an Implicit Function

AAAI 2022technical

Recovering detailed facial geometry from a set of calibrated multi-view images is valuable for its wide range of applications. Traditional multi-view stereo (MVS) methods adopt an optimization-based scheme to regularize the matching cost. Recently, learning-based methods integrate all these into an…

2022

Don’t Copy the Teacher: Data and Model Challenges in Embodied Dialogue

EMNLP 2022main

Embodied dialogue instruction following requires an agent to complete a complex sequence of tasks from a natural language exchange. The recent introduction of benchmarks raises the question of how best to train and evaluate models for this multi-turn, multi-agent, long-horizon task. This paper contr…

2020

AOT: Appearance Optimal Transport Based Identity Swapping for Forgery Detection

NeurIPS 2020poster

Recent studies have shown that the performance of forgery detection can be improved with diverse and challenging Deepfakes datasets. However, due to the lack of Deepfakes datasets with large variance in appearance, which can be hardly produced by recent identity swapping methods, the detection algor…

2020

Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning

IJCAI 2020poster

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on either disentangling the information in a single image or l…

Cited by 0SourcePDFScholar
2020

FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction

CVPR 2020poster

In this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific ex…

Cited by 361PDFcodeScholar
2020

Multi-agent Trajectory Prediction with Fuzzy Query Attention

NeurIPS 2020poster

Trajectory prediction for scenes with multiple agents and entities is a challenging problem in numerous domains such as traffic prediction, pedestrian tracking and path planning. We present a general architecture to address this challenge which models the crucial inductive biases of motion, namely,…

2020

SAPIEN: A SimulAted Part-Based Interactive ENvironment

CVPR 2020oral

Building home assistant robots has long been a goal for vision and robotics researchers. To achieve this task, a simulated environment with physically realistic simulation, sufficient articulated objects, and transferability to the real robot is indispensable. Existing environments achieve these req…

Cited by 560PDFcodeScholar
2020

Self-Supervised Human Depth Estimation From Monocular Videos

CVPR 2020poster

Previous methods on estimating detailed human depth often require supervised training with 'ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of…

Cited by 35PDFScholar
2019

CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark

CVPR 2019oral

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains challenging and inevitable in many scenarios. Moreover, current benchm…

Cited by 699PDFcodeScholar
2019

Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh Deformation

CVPR 2019oral

This paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template…

Cited by 173PDFcodeScholar
2019

S4G: Amodal Single-view Single-Shot SE(3) Grasp Detection in Cluttered Scenes

CoRL 2019

Grasping is among the most fundamental and long-lasting problems in robotics study. This paper studies the problem of 6-DoF(degree of freedom) grasping by a parallel gripper in a cluttered scene captured using a commodity depth sensor from a single viewpoint. We address the problem in a learning-bas