← Search

Mennatullah Siam

10 accepted papers

2026

Segmentation From Attention: Training-Free Layer Selection and One-Shot Tuning for Segmentation in VLMs

ICML 2026poster

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and image regions. This emergent ability enables zero-shot object detection and segmenta…

Cited by 0SourceScholar
2024

Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach

CVPR 2024poster

The emergence of attention-based transformer models has led to their extensive use in various tasks due to their superior generalization and transfer properties. Recent research has demonstrated that such models when prompted appropriately are excellent for few-shot inference. However such technique…

Cited by 9SourcePDFScholar
2023

AI-Assisted Tool for Early Diagnosis and Prevention of Colorectal Cancer in Africa

IJCAI 2023poster

Colorectal cancer (CRC) is considered the third most common cancer worldwide and is recently increasing in Africa. It is mostly diagnosed at an advanced state causing high fatality rates, which highlights the importance of CRC early diagnosis. There are various methods used to enable early diagnosis…

Cited by 0SourcePDFScholar
2023

MED-VT: Multiscale Encoder-Decoder Video Transformer With Application To Object Segmentation

CVPR 2023poster

Multiscale video transformers have been explored in a wide variety of vision tasks. To date, however, the multiscale processing has been confined to the encoder or decoder alone. We present a unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in videos. Multisca…

Cited by 22SourcePDFScholar
2022

A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information

CVPR 2022poster

Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been obser…

Cited by 22PDFcodeScholar
2020

Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings

IJCAI 2020poster

Significant progress has been made recently in developing few-shot object segmentation methods. Learning is shown to be successful in few-shot segmentation settings, using pixel-level, scribbles and bounding box supervision. This paper takes another approach, i.e., only requiring image-level label f…

Cited by 0SourcePDFScholar
2019

Online Object and Task Learning via Human Robot Interaction

ICRA 2019poster

This work describes the development of a robotic system that acquires knowledge incrementally through human interaction where new objects and motions are taught on the fly. The robotic system developed was one of the five finalists in the KUKA Innovation Award competition and demonstrated during the…

Cited by 32SourceScholar
2019

Video Object Segmentation using Teacher-Student Adaptation in a Human Robot Interaction (HRI) Setting

ICRA 2019poster

Video object segmentation is an essential task in robot manipulation to facilitate grasping and learning affordances. Incremental learning is important for robotics in unstructured environments. Inspired by the children learning process, human robot interaction (HRI) can be utilized to teach robots…

Cited by 108SourcecodeScholar
2018

Real-Time Segmentation with Appearance, Motion and Geometry

IROS 2018poster

Real-time Segmentation is of crucial importance to robotics related applications such as autonomous driving, driving assisted systems, and traffic monitoring from unmanned aerial vehicles imagery. We propose a novel two-stream convolutional network for motion segmentation, which exploits flow and ge…

Cited by 19SourceScholar