← Search

Quoc-Huy Tran

15 accepted papers

2025

Joint Self-Supervised Video Alignment and Action Segmentation

ICCV 2025poster

We introduce a novel approach for simultaneous self-supervised video alignment and action segmentation based on a unified optimal transport framework. In particular, we first tackle self-supervised video alignment by developing a fused Gromov-Wasserstein optimal transport formulation with a structur…

Cited by 0SourcePDFScholar
2024

Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

ICRA 2024poster

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs and apply Graph Convolutional Networks (GCNs) for spatiotempor…

Cited by 5SourceScholar
2022

Inductive and Transductive Few-Shot Video Classification via Appearance and Temporal Alignments

ECCV 2022poster

"We present a novel method for few-shot video classification, which performs appearance and temporal alignments. In particular, given a pair of query and support videos, we conduct appearance alignment via frame-level feature matching to achieve the appearance similarity score between the videos, wh…

2022

Timestamp-Supervised Action Segmentation with Graph Convolutional Networks

IROS 2022poster

We introduce a novel approach for temporal activity segmentation with timestamp supervision. Our main contribution is a graph convolutional network, which is learned in an end-to-end manner to exploit both frame features and connections between neighboring frames to generate dense framewise labels f…

Cited by 18SourceScholar
2022

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

CVPR 2022poster

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequ…

Cited by 70PDFcodeScholar
2021

Learning by Aligning Videos in Time

CVPR 2021poster

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment loss and temporal regularization terms, which can be used as…

Cited by 86PDFScholar
2021

POODLE: Improving Few-shot Learning via Penalizing Out-of-Distribution Samples

NeurIPS 2021poster

In this work, we propose to use out-of-distribution samples, i.e., unlabeled samples coming from outside the target classes, to improve few-shot learning. Specifically, we exploit the easily available out-of-distribution samples to drive the classifier to avoid irrelevant features by maximizing the…

2020

Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling

ECCV 2020poster

Monocular visual odometry (VO) suffers severely from error accumulation during frame-to-frame pose estimation. In this paper, we present a self-supervised learning method for VO with special consideration for consistency over longer sequences. To this end, we model the long-term dependency in pose p…

2020

Pseudo RGB-D for Self-Improving Monocular SLAM and Depth Prediction

ECCV 2020poster

Classical monocular Simultaneous Localization And Mapping (SLAM) and the recently emerging convolutional neural networks (CNNs) for monocular depth prediction represent two largely disjoint approaches towards building a 3D map of the surrounding environment. In this paper, we demonstrate that the co…

2019

Degeneracy in Self-Calibration Revisited and a Deep Learning Solution for Uncalibrated SLAM

IROS 2019poster

Self-calibration of camera intrinsics and radial distortion has a long history of research in the computer vision community. However, it remains rare to see real applications of such techniques to modern Simultaneous Localization And Mapping (SLAM) systems, especially in driving scenarios. In this p…

Cited by 27SourceScholar
2019

Learning Structure-And-Motion-Aware Rolling Shutter Correction

CVPR 2019oral

An exact method of correcting the rolling shutter (RS) effect requires recovering the underlying geometry, i.e. the scene structures and the camera motions between scanlines or between views. However, the multiple-view geometry for RS cameras is much more complicated than its global shutter (GS) cou…

Cited by 64PDFScholar
2018

Hierarchical Metric Learning and Matching for 2D and 3D Geometric Correspondences

ECCV 2018poster

Interest point descriptors have fueled progress on almost every problem in computer vision. Recent advances in deep neural networks have enabled task-specific learned descriptors that outperform hand-crafted descriptors on many problems. We demonstrate that commonly used metric learning approaches d…

Cited by 56SourcePDFScholar
2017

Deep Supervision With Shape Concepts for Occlusion-Aware 3D Object Parsing

CVPR 2017poster

Monocular 3D object parsing is highly desirable in various scenarios including occlusion reasoning and holistic scene interpretation. We present a deep convolutional neural network (CNN) architecture to localize semantic parts in 2D image and 3D space while inferring their visibility states, given a…

Cited by 110PDFScholar
2016

A Continuous Occlusion Model for Road Scene Understanding

CVPR 2016poster

We present a physically interpretable, continuous 3D model for handling occlusions with applications to road scene understanding. We probabilistically assign each point in space to an object with a theoretical modeling of the reflection and transmission probabilities for the corresponding camera ray…

Cited by 37PDFScholar