← Search

Ruoshi Liu

15 accepted papers

2024

Controlling the World by Sleight of Hand

ECCV 2024oral

"Humans naturally build mental models of object interactions and dynamics, allowing them to imagine how their surroundings will change if they take a certain action. While generative models today have shown impressive results on generating/editing images unconditionally or conditioned on text, curre…

Cited by 3SourcePDFScholar
2024

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

CoRL 2024poster

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a vi…

Cited by 26SourceScholar
2024

EraseDraw : Learning to Insert Objects by Erasing Them from Images

ECCV 2024poster

"Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often fail by making global changes to the image, inserting objects in unrealistic spatial locations, and generating inaccurat…

Cited by 2SourcePDFScholar
2024

GES : Generalized Exponential Splatting for Efficient Radiance Field Rendering

CVPR 2024poster

Advancements in 3D Gaussian Splatting have significantly accelerated 3D reconstruction and generation. However it may require a large number of Gaussians which creates a substantial memory footprint. This paper introduces GES (Generalized Exponential Splatting) a novel representation that employs Ge…

2024

Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

ECCV 2024oral

"Accurate reconstruction of complex dynamic scenes from just a single viewpoint continues to be a challenging task in computer vision. Current dynamic novel view synthesis methods typically require videos from many different camera viewpoints, necessitating careful recording setups, and significantl…

Cited by 23SourcePDFScholar
2024

Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape

ICLR 2024poster

Synthesizing novel 3D models that resemble the input example as long been pursued by graphics artists and machine learning researchers. In this paper, we present Sin3DM, a diffusion model that learns the internal patch distribution from a single 3D textured shape and generates high-quality variation…

2024

pix2gestalt: Amodal Segmentation by Synthesizing Wholes

CVPR 2024highlight

We introduce pix2gestalt a framework for zero-shot amodal segmentation which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task we learn a…

2023

Objaverse-XL: A Universe of 10M+ 3D Objects

NeurIPS 2023poster

Natural language processing and 2D vision models have attained remarkable proficiency on many tasks primarily by escalating the scale of training data. However, 3D vision tasks have not seen the same progress, in part due to the challenges of acquiring high-quality 3D data. In this work, we present…

Cited by 393SourcePDFScholar
2023

What You Can Reconstruct From a Shadow

CVPR 2023poster

3D reconstruction is a fundamental problem in computer vision, and the task is especially challenging when the object to reconstruct is partially or fully occluded. We introduce a method that uses the shadows cast by an unobserved object in order to infer the possible 3D volumes under occlusion. We…

Cited by 3SourcePDFScholar
2023

Zero-1-to-3: Zero-shot One Image to 3D Object

ICCV 2023poster

We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this underconstrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional di…

Cited by 1020PDFcodeScholar