← Search

Hsiao-Yu Tung

13 accepted papers

2023

3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes

NeurIPS 2023poster

Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on ex…

Cited by 8SourcePDFScholar
2023

FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

ICLR 2023top-25%

Humans manipulate various kinds of fluids in their everyday life: creating latte art, scooping floating objects from water, rolling an ice cream cone, etc. Using robots to augment or replace human labors in these daily settings remain as a challenging task due to the multifaceted complexities of flu…

2023

H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from Interactions

ICRA 2023poster

The world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in…

Cited by 3SourceScholar
2023

Tactile-Filter: Interactive Tactile Perception for Part Mating

RSS 2023poster

Humans rely on touch and tactile sensing for a lot of dexterous manipulation tasks. Our tactile sensing provides us with a lot of information regarding contact formations as well as geometric information about objects during any interaction. With this motivation, vision-based tactile sensors are bei…

2021

Disentangling 3D Prototypical Networks for Few-Shot Concept Learning

ICLR 2021poster

We present neural architectures that disentangle RGB-D images into objects’ shapes and styles and a map of the background scene, and explore their applications for few-shot 3D object detection and few-shot concept classification. Our networks incorporate architectural biases that reflect the image f…

2021

HyperDynamics: Meta-Learning Object and Agent Dynamics with Hypernetworks

ICLR 2021poster

We propose HyperDynamics, a dynamics meta-learning framework that conditions on an agent’s interactions with the environment and optionally its visual observations, and generates the parameters of neural dynamics models based on inferred properties of the dynamical system. Physical and visual proper…

Cited by 27SourcePDFScholar
2021

Physion: Evaluating Physical Prediction from Vision in Humans and Machines

NeurIPS 2021poster

While current vision algorithms excel at many challenging tasks, it is unclear how well they understand the physical dynamics of real-world environments. Here we introduce Physion, a dataset and benchmark for rigorously evaluating the ability to predict how physical scenarios will evolve over time.…

Cited by 82SourcecodeScholar
2021

Visually-Grounded Library of Behaviors for Manipulating Diverse Objects across Diverse Configurations and Views

CoRL 2021poster

We propose a visually-grounded library of behaviors approach for learning to manipulate diverse objects across varying initial and goal configurations and camera placements. Our key innovation is to disentangle the standard image-to-action mapping into two separate modules that use different types o…

Cited by 1SourceScholar
2020

3D-OES: Viewpoint-Invariant Object-Factorized Environment Simulators

CoRL 2020

We propose an action-conditioned dynamics model that predicts scene changes caused by object and agent interactions in a viewpoint-invariant 3D neural scene representation space, inferred from RGB-D videos. In this 3D feature space, objects do not interfere with one another and their appearance pers

Cited by 0SourcePDFScholar
2018

Reward Learning From Narrated Demonstrations

CVPR 2018poster

Humans effortlessly “program” one another by communicating goals and desires in natural language. In contrast, humans program robotic behaviours by indicating desired object locations and poses to be achieved [5], by providing RGB images of goal configurations [19], or supplying a demonstration to b…

Cited by 46SourcePDFScholar
2017

Generative Models and Model Criticism via Optimized Maximum Mean Discrepancy

ICLR 2017poster

We propose a method to optimize the representation and distinguishability of samples from two probability distributions, by maximizing the estimated power of a statistical test based on the maximum mean discrepancy (MMD). This optimized MMD is applied to the setting of unsupervised learning by gener…

Cited by 252SourcecodeScholar
2017

Self-supervised Learning of Motion Capture

NeurIPS 2017spotlight

Current state-of-the-art solutions for motion capture from a single camera are optimization driven: they optimize the parameters of a 3D human model so that its re-projection matches measurements in the video (e.g. person segmentation, optical flow, keypoint detections etc.). Optimization models are…

Cited by 369SourcePDFScholar
2015

Fast and Guaranteed Tensor Decomposition via Sketching

NeurIPS 2015spotlight

Tensor CANDECOMP/PARAFAC (CP) decomposition has wide applications in statistical learning of latent variable models and in data mining. In this paper, we propose fast and randomized tensor CP decomposition algorithms based on sketching. We build on the idea of count sketches, but introduce many nove…

Cited by 160SourcePDFScholar