← Search

Winston Hsu

4 accepted papers

2021

OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding

NAACL 2021long

To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded objects. However, current VG works are limited in working environments, such as offices and warehouses, where objects ar…

2020

GDN: A Coarse-To-Fine (C2F) Representation for End-To-End 6-DoF Grasp Detection

CoRL 2020

We proposed an end-to-end grasp detection network, Grasp Detection Network (GDN), cooperated with a novel coarse-to-fine (C2F) grasp representation design to detect diverse and accurate 6-DoF grasps based on point clouds. Compared to previous two-stage approaches which sample and evaluate multiple g

Cited by 0SourcePDFScholar
2019

Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGAN

ICCV 2019poster

Free-form video inpainting is a very challenging task that could be widely used for video editing such as text removal. Existing patch-based methods could not handle non-repetitive structures such as faces, while directly applying image-based inpainting models to videos will result in temporal incon…

Cited by 270PDFcodeScholar
2017

Joint Sequence Learning and Cross-Modality Convolution for 3D Biomedical Segmentation

CVPR 2017poster

Deep learning models such as convolutional neural network have been widely used in 3D biomedical segmentation and achieve state-of-the-art performance. However, most of them often adapt a single modality or stack multiple modalities as different input channels, which ignores the correlations among t…

Cited by 226PDFScholar