← Search

Yao-Hung Hubert Tsai

20 accepted papers

2024

KPConvX: Modernizing Kernel Point Convolution with Kernel Attention

CVPR 2024poster

In the field of deep point cloud understanding KPConv is a unique architecture that uses kernel points to locate convolutional weights in space instead of relying on Multi-Layer Perceptron (MLP) encodings. While it initially achieved success it has since been surpassed by recent MLP networks that em…

2023

Self-Supervised Object Goal Navigation with In-Situ Finetuning

IROS 2023poster

A household robot should be able to navigate to target objects without requiring users to first annotate everything in their home. Most current approaches to object navigation do not test on real robots and rely solely on reconstructed scans of houses and their expensively labeled semantic 3D meshes…

Cited by 7SourceScholar
2022

Conditional Contrastive Learning with Kernel

ICLR 2022poster

Conditional contrastive learning frameworks consider the conditional sampling procedure that constructs positive or negative data pairs conditioned on specific variables. Fair contrastive learning constructs negative pairs, for example, from the same gender (conditioning on sensitive information), w…

2022

Greedy modality selection via approximate submodular maximization

UAI 2022poster

Multimodal learning considers learning from multi-modality data, aiming to fuse heterogeneous sources of information. However, it is not always feasible to leverage all available modalities due to memory constraints. Further, training on all the modalities may be inefficient when redundant informati…

Cited by 4SourcePDFScholar
2022

Learning Weakly-supervised Contrastive Representations

ICLR 2022poster

We argue that a form of the valuable information provided by the auxiliary information is its implied data clustering information. For instance, considering hashtags as auxiliary information, we can hypothesize that an Instagram image will be semantically more similar with the same hashtags. With th…

2022

Paraphrasing Is All You Need for Novel Object Captioning

NeurIPS 2022accept

Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we pr…

Cited by 5SourcePDFScholar
2021

Hubert: How Much Can a Bad Teacher Benefit ASR Pre-Training?

ICASSP 2021accepted

Compared to vision and language applications, self-supervised pre-training approaches for ASR are challenged by three unique problems: (1) There are multiple sound units in each input utterance, (2) With audio-only pre-training, there is no lexicon of sound units, and (3) Sound units have variable l…

Cited by 0SourceScholar
2021

Self-supervised Learning from a Multi-view Perspective

ICLR 2021poster

As a subset of unsupervised representation learning, self-supervised representation learning adopts self-defined signals as supervision and uses the learned representation for downstream tasks, such as object detection and image captioning. Many proposed approaches for self-supervised learning follo…

2021

Self-supervised Representation Learning with Relative Predictive Coding

ICLR 2021poster

This paper introduces Relative Predictive Coding (RPC), a new contrastive representation learning objective that maintains a good balance among training stability, minibatch size sensitivity, and downstream task performance. The key to the success of RPC is two-fold. First, RPC introduces the relati…

2020

Capsules with Inverted Dot-Product Attention Routing

ICLR 2020poster

We introduce a new routing algorithm for capsule networks, in which a child capsule is routed to a parent based only on agreement between the parent's state and the child's vote. The new mechanism 1) designs routing via inverted dot-product attention; 2) imposes Layer Normalization as normalization…

Cited by 115SourceScholar
2020

Complex Transformer: A Framework for Modeling Complex-Valued Sequence

ICASSP 2020accepted

While deep learning has received a surge of interest in a variety of fields in recent years, major deep learning models barely use complex numbers. However, speech, signal and audio data are naturally complex-valued after Fourier Transform, and studies have shown a potentially richer representation…

Cited by 0SourceScholar
2020

Neural Methods for Point-wise Dependency Estimation

NeurIPS 2020spotlight

Since its inception, the neural estimation of mutual information (MI) has demonstrated the empirical success of modeling expected dependency between high-dimensional random variables. However, MI is an aggregate statistic and cannot be used to measure point-wise dependency between different events.…

2019

Learning Factorized Multimodal Representations

ICLR 2019poster

Learning multimodal representations is a fundamentally complex research problem due to the presence of multiple heterogeneous sources of information. Although the presence of multiple modalities provides additional valuable information, there are two key challenges to address when learning from mult…

2019

Learning Neural Networks with Adaptive Regularization

NeurIPS 2019poster

Feed-forward neural networks can be understood as a combination of an intermediate representation and a linear hypothesis. While most previous works aim to diversify the representations, we explore the complementary direction by performing an adaptive and data-dependent regularization motivated by t…

2019

Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator

ICLR 2019poster

Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data, designing divergence measure that is interpretable and can ha…

Cited by 28SourcePDFScholar
2019

Video Relationship Reasoning Using Gated Spatio-Temporal Energy Graph

CVPR 2019poster

Visual relationship reasoning is a crucial yet challenging task for understanding rich interactions across visual concepts. For example, a relationship \ man, open, door\ involves a complex relation \ open\ between concrete entities \ man, door\ . While much of the existing work has studied this p…

Cited by 127PDFcodeScholar
2016

Heterogeneous domain adaptation with label and structure consistency

ICASSP 2016accepted

Domain adaptation is a challenging task, since it associates data collected from different domains or exhibiting distinct distributions. In this paper, we particularly focus on adapting cross-domain data with distinct feature dimensions or representations. Thus, this is referred to as the task of he…

Cited by 0SourceScholar
2016

Learning Cross-Domain Landmarks for Heterogeneous Domain Adaptation

CVPR 2016poster

While domain adaptation (DA) aims to associate the learning tasks across data domains, heterogeneous domain adaptation (HDA) particularly deals with learning from cross-domain data which are of different types of features. In other words, for HDA, data from source and target domains are observed in…

Cited by 244PDFScholar
2015

Unsupervised Domain Adaptation With Imbalanced Cross-Domain Data

ICCV 2015poster

We address a challenging unsupervised domain adaptation problem with imbalanced cross-domain data. For standard unsupervised domain adaptation, one typically obtains labeled data in the source domain and only observes unlabeled data in the target domain. However, most existing works do not consider…

Cited by 90PDFScholar