← Search

Mohammed Bennamoun

44 accepted papers

2026

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

CVPR 2026

Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language Models (MLLMs) typically compress an entire WSI into a single

Cited by 0SourcecodeScholar
2026

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

ICML 2026poster

Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, where only a single labeled exemplar is available for each novel action. A key challenge is learning representations that capture the hierarchical and…

Cited by 0SourceScholar
2026

SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

CVPR 2026

Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typically align skeleton features with textual embeddings within a shared latent space. However, the absence of contextual

Cited by 0SourcecodeScholar
2025

Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

NeurIPS 2025poster

We introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-…

Cited by 0SourcecodeScholar
2025

Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis

CVPR 2025poster

We propose a novel framework for the statistical analysis of genus-zero 4D surfaces, i.e., 3D surfaces that deform and evolve overtime. This problem is particularly challenging due to the arbitrary parameterizations of these surfaces and their varying deformation speeds, necessitating effective spat…

Cited by 0SourcePDFScholar
2025

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

ICML 2025poster

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS)…

2025

Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation

CVPR 2025poster

In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tis…

2025

STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

CVPR 2025highlight

Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a c…

2025

Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM

NeurIPS 2025poster

Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like “A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding” requires simultaneous proc…

Cited by 0SourceScholar
2024

A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures

ECCV 2024oral

"We propose the first comprehensive approach for modeling and analyzing the spatiotemporal shape variability in tree-like 4D objects, 3D objects whose shapes bend, stretch and change in their branching structure over time as they deform, grow, and interact with their environment. Our key contributio…

2024

AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion

ECCV 2024poster

"Point cloud completion involves inferring missing parts of 3D objects from incomplete point cloud data. It requires a model that understands the global structure of the object and reconstructs local details. To this end, we propose a global perception and local attention network, termed AEDNet, for…

Cited by 1SourcePDFScholar
2024

CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment

CVPR 2024poster

This paper proposes Comprehensive Pathology Language Image Pre-training (CPLIP) a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision language models by leveraging extens…

2024

DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition

ECCV 2024poster

"Neuromorphic sensors, specifically event cameras, revolutionize visual data acquisition by capturing pixel intensity changes with exceptional dynamic range, minimal latency, and energy efficiency, setting them apart from conventional frame-based cameras. The distinctive capabilities of event camera…

2024

Language Model Guided Interpretable Video Action Reasoning

CVPR 2024poster

Although neural networks excel in video action recognition tasks their "black-box" nature makes it challenging to understand the rationale behind their decisions. Recent approaches used inherently interpretable models to analyze video actions in a manner akin to human reasoning. However it has been…

2024

Referring Human Pose and Mask Estimation In the Wild

NeurIPS 2024poster

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to pr…

2023

B-Pose: Bayesian Deep Network for Camera 6-DoF Pose Estimation From RGB Images

RA-L 2023

Camera pose estimation has long relied on geometry-based approaches and sparse 2D-3D keypoint correspondences. With the advent of deep learning methods, the estimation of camera pose parameters, i.e., the six parameters that describe position and rotation denoted by 6 Degrees of Freedom (6-DoF), has

Cited by 9SourceScholar
2023

Extended Expectation Maximization for Under-Fitted Models

ICASSP 2023accepted

In this paper, we generalize the well-known Expectation Maximization (EM) algorithm using the α−divergence for Gaussian Mixture Model (GMM). This approach is used in robust subspace detection when the number of parameters is kept small to avoid overfitting and large estimation variances. The level o…

Cited by 0SourceScholar
2023

Learning Multi-Modal Class-Specific Tokens for Weakly Supervised Dense Object Localization

CVPR 2023poster

Weakly supervised dense object localization (WSDOL) relies generally on Class Activation Mapping (CAM), which exploits the correlation between the class weights of the image classifier and the pixel-level features. Due to the limited ability to address intra-class variations, the image classifier ca…

2023

Reinforced Learning for Label-Efficient 3D Face Reconstruction

ICRA 2023poster

3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face im…

Cited by 1SourceScholar
2023

Spectrum-guided Multi-granularity Referring Video Object Segmentation

ICCV 2023poster

Current referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perc…

Cited by 54PDFcodeScholar
2023

UE4-NeRF:Neural Radiance Field for Real-Time Rendering of Large-Scale Scene

NeurIPS 2023poster

Neural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real…

2023

VAPCNet: Viewpoint-Aware 3D Point Cloud Completion

ICCV 2023poster

Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating t…

Cited by 12PDFcodeScholar
2022

Active-Passive SimStereo - Benchmarking the Cross-Generalization Capabilities of Deep Learning-based Stereo Methods

NeurIPS 2022accept

In stereo vision, self-similar or bland regions can make it difficult to match patches between two images. Active stereo-based methods mitigate this problem by projecting a pseudo-random pattern on the scene so that each patch of an image pair can be identified without ambiguity. However, the projec…

Cited by 4SourcePDFScholar
2022

Adversary Distillation for One-Shot Attacks on 3D Target Tracking

ICASSP 2022accepted

Considering the vulnerability of existing deep models in the adversarial scenario, the robustness of 3D target tracking is not guaranteed. In this paper, we present an efficient generation based adversarial attack, termed Adversary Distillation Network (AD-Net), which is able to distract a victim tr…

Cited by 0SourceScholar
2022

Multi-Class Token Transformer for Weakly Supervised Semantic Segmentation

CVPR 2022poster

This paper proposes a new transformer-based framework to learn class-specific object localization maps as pseudo labels for weakly supervised semantic segmentation (WSSS). Inspired by the fact that the attended regions of the one-class token in the standard vision transformer can be leveraged to for…

Cited by 301PDFcodeScholar
2021

CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image Saliency

CVPR 2021poster

Backpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, class-insensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, res…

Cited by 69PDFcodeScholar
2021

Leveraging Auxiliary Tasks With Affinity Learning for Weakly Supervised Semantic Segmentation

ICCV 2021poster

Semantic segmentation is a challenging task in the absence of densely labelled data. Only relying on class activation maps (CAM) with image-level labels provides deficient segmentation supervision. Prior works thus consider pre-trained models to produce coarse saliency maps to guide the generation o…

Cited by 154PDFcodeScholar
2020

Efficient Scene Text Detection with Textual Attention Tower

ICASSP 2020accepted

Scene text detection has received attention for years and achieved an impressive performance across various benchmarks. In this work, we propose an efficient and accurate approach to detect multi-oriented text in scene images. The proposed feature fusion mechanism allows us to use a shallower networ…

Cited by 0SourceScholar
2019

An Improved Approach to Weakly Supervised Semantic Segmentation

ICASSP 2019accepted

Weakly supervised semantic segmentation with image-level labels is of great significance since it alleviates the dependency on dense annotations. However, it is a challenging task as it aims to achieve a mapping from high-level semantics to low-level features. In this work, we propose a three-step m…

Cited by 0SourceScholar
2018

Attention in Convolutional LSTM for Gesture Recognition

NeurIPS 2018poster

Convolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the…

2018

Classification of Corals in Reflectance and Fluorescence Images Using Convolutional Neural Network Representations

ICASSP 2018accepted

Coral species, with complex morphology and ambiguous boundaries, pose a great challenge for automated classification. CNN activations, which are extracted from fully connected layers of deep networks (FC features), have been successfully used as powerful universal representations in many visual task…

Cited by 0SourceScholar
2018

NNEval: Neural Network based Evaluation Metric for Image Captioning

ECCV 2018poster

The automatic evaluation of image descriptions is an intricate task, and it is highly important in the development and fine-grained analysis of captioning systems. Existing metrics to automatically evaluate image captioning systems fail to achieve a satisfactory level of correlation with human judge…

Cited by 26SourcePDFScholar
2017

A New Representation of Skeleton Sequences for 3D Action Recognition

CVPR 2017poster

This paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several frames for spatial temporal feature learning using deep neural…

Cited by 1080PDFScholar
2015

Contractive Rectifier Networks for Nonlinear Maximum Margin Classification

ICCV 2015poster

To find the optimal nonlinear separating boundary with maximum margin in the input data space, this paper proposes Contractive Rectifier Networks (CRNs), wherein the hidden-layer transformations are restricted to be contraction mappings. The contractive constraints ensure that the achieved separatin…

Cited by 13PDFScholar
2015

Efficient RGB-D object categorization using cascaded ensembles of randomized decision trees

ICRA 2015poster

This paper presents an efficient framework for the categorization of objects in real-world scenes (captured with an RGB-D sensor). The proposed framework uses ensembles of randomized decision trees in a hierarchical cascaded architecture to compute consistent object-class inferences of unseen object…

Cited by 34SourceScholar
2015

How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances?

ICML 2015poster

This paper investigates how hidden layers of deep rectifier networks are capable of transforming two or more pattern sets to be linearly separable while preserving the distances with a guaranteed degree, and proves the universal classification power of such distance preserving rectifier networks. Th…

Cited by 34SourcePDFScholar
2015

Listening With Your Eyes: Towards a Practical Visual Speech Recognition System Using Deep Boltzmann Machines

ICCV 2015poster

This paper presents a novel feature learning method for visual speech recognition using Deep Boltzmann Machines (DBM). Unlike all existing visual feature extraction techniques which solely extracts features from video sequences, our method is able to explore both acoustic information and visual info…

Cited by 49PDFScholar
2015

Separating Objects and Clutter in Indoor Scenes

CVPR 2015poster

Objects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we…

Cited by 26SourcePDFScholar