← Search

Andrey Konin

6 accepted papers

2025

Joint Self-Supervised Video Alignment and Action Segmentation

ICCV 2025poster

We introduce a novel approach for simultaneous self-supervised video alignment and action segmentation based on a unified optimal transport framework. In particular, we first tackle self-supervised video alignment by developing a fused Gromov-Wasserstein optimal transport formulation with a structur…

Cited by 0SourcePDFScholar
2024

Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

ICRA 2024poster

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs and apply Graph Convolutional Networks (GCNs) for spatiotempor…

Cited by 5SourceScholar
2024

Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion

ECCV 2024poster

"This paper presents a self-supervised temporal video alignment framework which is useful for several fine-grained human activity understanding applications. In contrast with the state-of-the-art method of CASA, where sequences of 3D skeleton coordinates are taken directly as input, our key idea is…

Cited by 3SourcePDFScholar
2022

Timestamp-Supervised Action Segmentation with Graph Convolutional Networks

IROS 2022poster

We introduce a novel approach for temporal activity segmentation with timestamp supervision. Our main contribution is a graph convolutional network, which is learned in an end-to-end manner to exploit both frame features and connections between neighboring frames to generate dense framewise labels f…

Cited by 18SourceScholar
2022

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

CVPR 2022poster

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequ…

Cited by 70PDFcodeScholar
2021

Learning by Aligning Videos in Time

CVPR 2021poster

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment loss and temporal regularization terms, which can be used as…

Cited by 86PDFScholar