← Search

Ali Athar

7 accepted papers

2025

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

NeurIPS 2025poster

This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome limitations in existing image-text datasets that often lack detailed, scene-com…

Cited by 0SourceScholar
2025

ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

CVPR 2025poster

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically invo…

2023

TarViS: A Unified Approach for Target-Based Video Segmentation

CVPR 2023highlight

The general domain of video segmentation is currently fragmented into different tasks spanning multiple benchmarks. Despite rapid progress in the state-of-the-art, current methods are overwhelmingly task-specific and cannot conceptually generalize to other tasks. Inspired by recent approaches with m…

2022

HODOR: High-Level Object Descriptors for Object Re-Segmentation in Video Learned From Static Images

CVPR 2022oral

Existing state-of-the-art methods for Video Object Segmentation (VOS) learn low-level pixel-to-pixel correspondences between frames to propagate object masks across video. This requires a large amount of densely annotated video data, which is costly to annotate, and largely redundant since frames wi…

Cited by 30PDFcodeScholar
2020

STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos

ECCV 2020poster

Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in individual frames, and then associate these detections over time…

2016

Whole-body motion planning for humanoid robots with heuristic search

IROS 2016poster

The task of whole-body motion planning for humanoid robots is challenging due to its high-DOF nature, stability constraints, and the need for obstacle avoidance and movements that are efficient. Over the years, various approaches have been adopted to solve this problem such as bounding-box models an…

Cited by 5SourceScholar