← Search

M. Zeeshan Zia

9 accepted papers

2025

Joint Self-Supervised Video Alignment and Action Segmentation

ICCV 2025poster

We introduce a novel approach for simultaneous self-supervised video alignment and action segmentation based on a unified optimal transport framework. In particular, we first tackle self-supervised video alignment by developing a fused Gromov-Wasserstein optimal transport formulation with a structur…

Cited by 0SourcePDFScholar
2024

Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

ICRA 2024poster

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs and apply Graph Convolutional Networks (GCNs) for spatiotempor…

Cited by 5SourceScholar
2022

Timestamp-Supervised Action Segmentation with Graph Convolutional Networks

IROS 2022poster

We introduce a novel approach for temporal activity segmentation with timestamp supervision. Our main contribution is a graph convolutional network, which is learned in an end-to-end manner to exploit both frame features and connections between neighboring frames to generate dense framewise labels f…

Cited by 18SourceScholar
2022

Unsupervised Action Segmentation by Joint Representation Learning and Online Clustering

CVPR 2022poster

We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequ…

Cited by 70PDFcodeScholar
2018

Hierarchical Metric Learning and Matching for 2D and 3D Geometric Correspondences

ECCV 2018poster

Interest point descriptors have fueled progress on almost every problem in computer vision. Recent advances in deep neural networks have enabled task-specific learned descriptors that outperform hand-crafted descriptors on many problems. We demonstrate that commonly used metric learning approaches d…

Cited by 56SourcePDFScholar
2017

Deep Supervision With Shape Concepts for Occlusion-Aware 3D Object Parsing

CVPR 2017poster

Monocular 3D object parsing is highly desirable in various scenarios including occlusion reasoning and holistic scene interpretation. We present a deep convolutional neural network (CNN) architecture to localize semantic parts in 2D image and 3D space while inferring their visibility states, given a…

Cited by 110PDFScholar
2016

Comparative design space exploration of dense and semi-dense SLAM

ICRA 2016

SLAM has matured significantly over the past few years, and is beginning to appear in serious commercial products. While new SLAM systems are being proposed at every conference, evaluation is often restricted to qualitative visualizations or accuracy estimation against a ground truth. This is due to

Cited by 26SourceScholar
2016

Monocular reconstruction of vehicles: Combining SLAM with shape priors

ICRA 2016

Reasoning about objects in images and videos using 3D representations is re-emerging as a popular paradigm in computer vision. Specifically, in the context of scene understanding for roads, 3D vehicle detection and tracking from monocular videos still needs a lot of attention to enable practical app

Cited by 50SourceScholar
2015

Introducing SLAMBench, a performance and accuracy benchmarking methodology for SLAM

ICRA 2015poster

Real-time dense computer vision and SLAM offer great potential for a new level of scene modelling, tracking and real environmental interaction for many types of robot, but their high computational requirements mean that use on mass market embedded platforms is challenging. Meanwhile, trends in low-c…

Cited by 211SourceScholar