← Search

Djamila Aouada

26 accepted papers

2026

Annotation Free Spacecraft Detection and Segmentation Using Vision Language Models

ICRA 2026poster

Vision Language Models (VLMs) have demonstrated remarkable performance in open-world zero-shot visual recognition. However, their potential in space-related applications remains largely unexplored. In the space domain, accurate manual annotation is particularly challenging due to factors such as low…

2026

Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation

CVPR 2026

In this paper, we address the problem of 6-DoF object pose estimation from a single RGB image. Indirect methods that typically predict intermediate 2D keypoints, followed by a Perspective-n-Point solver, have shown great performance. Direct approaches, which regress the pose in an end-to-end manner,

Cited by 0SourceScholar
2025

CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers

ICCV 2025poster

We propose CAD-Assistant, a general-purpose CAD agent for AI-assisted design. Our approach is based on a powerful Vision and Large Language Model (VLLM) as a planner and a tool-augmentation paradigm using CAD-specific tools. CAD-Assistant addresses multimodal user queries by generating actions that…

2025

CAD-Recode: Reverse Engineering CAD Code from Point Clouds

ICCV 2025poster

Computer-Aided Design (CAD) models are typically constructed by sequentially drawing parametric sketches and applying CAD operations to obtain a 3D model. The problem of 3D CAD reverse engineering consists of reconstructing the sketch and CAD operation sequences from 3D representations such as point…

2025

Hybrid Attention for Robust RGB-T Pedestrian Detection in Real-World Conditions

RA-L 2025

Multispectral pedestrian detection has gained significant attention in recent years, particularly in autonomous driving applications. To address the challenges posed by adversarial illumination conditions, the combination of thermal and visible images has demonstrated its advantages. However, existi

Cited by 2SourceScholar
2025

MiCADangelo: Fine-Grained Reconstruction of Constrained CAD Models from 3D Scans

NeurIPS 2025poster

Computer-Aided Design (CAD) plays a foundational role in modern manufacturing and product development, often requiring designers to modify or build upon existing models. Converting 3D scans into parametric CAD representations—a process known as CAD reverse engineering—remains a significant challenge…

Cited by 0SourceScholar
2025

Uncertainty-Aware Knowledge Distillation for Compact and Efficient 6DoF Pose Estimation

IROS 2025

Compact and efficient 6DoF object pose estimation is crucial in applications such as robotics, augmented reality, and space autonomous navigation systems, where lightweight models are critical for real-time accurate performance. This paper introduces a novel uncertainty-aware end-to-end Knowledge Di

Cited by 2SourceScholar
2025

Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection

ICCV 2025poster

Detecting deepfake videos is highly challenging given the complexity of characterizing spatio-temporal artifacts. Most existing methods rely on binary classifiers trained using real and fake image sequences, therefore hindering their generalization capabilities to unseen generation methods. Moreover…

2024

A Hitchhiker's Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning

NeurIPS 2024poster

Explainability in artificial intelligence is crucial for restoring trust, particularly in areas like face forgery detection, where viewers often struggle to distinguish between real and fabricated content. Vision and Large Language Models (VLLM) bridge computer vision and natural language, offering…

2024

CAD-SIGNet: CAD Language Inference from Point Clouds using Layer-wise Sketch Instance Guided Attention

CVPR 2024highlight

Reverse engineering in the realm of Computer-Aided Design (CAD) has been a longstanding aspiration though not yet entirely realized. Its primary aim is to uncover the CAD process behind a physical object given its 3D scan. We propose CAD-SIGNet an end-to-end trainable and auto-regressive architectur…

Cited by 18SourcePDFScholar
2024

LAA-Net: Localized Artifact Attention Network for Quality-Agnostic and Generalizable Deepfake Detection

CVPR 2024poster

This paper introduces a novel approach for high-quality deepfake detection called Localized Artifact Attention Network (LAA-Net). Existing methods for high-quality deepfake detection are mainly based on a supervised binary classifier coupled with an implicit attention mechanism. As a result they do…

2024

SPADES: A Realistic Spacecraft Pose Estimation Dataset using Event Sensing

ICRA 2024poster

In recent years, there has been a growing demand for improved autonomy for in-orbit operations such as rendezvous, docking, and proximity manoeuvres, leading to increased interest in employing Deep Learning-based Spacecraft Pose Estimation techniques. However, due to limited access to real target da…

Cited by 6SourceScholar
2024

SpelsNet: Surface Primitive Elements Segmentation by B-Rep Graph Structure Supervision

NeurIPS 2024poster

Within the realm of Computer-Aided Design (CAD), Boundary-Representation (B-Rep) is the standard option for modeling shapes. We present SpelsNet, a neural architecture for the segmentation of 3D point clouds into surface primitive elements under topological supervision of its B-Rep graph structure.…

Cited by 1SourcePDFScholar
2024

TransCAD: A Hierarchical Transformer for CAD Sequence Inference from Point Clouds

ECCV 2024poster

"3D reverse engineering, in which a CAD model is inferred given a 3D scan of a physical object, is a research direction that offers many promising practical applications. This paper proposes , an end-to-end transformer-based architecture that predicts the CAD sequence from a point cloud. leverages t…

Cited by 40SourcePDFScholar
2023

3D-Aware Object Localization using Gaussian Implicit Occupancy Function

IROS 2023poster

To automatically localize a target object in an image is crucial for many computer vision applications. To represent the 2D object, ellipse labels have recently been identified as a promising alternative to axis-aligned bounding boxes. This paper further considers 3D-aware ellipse labels, i.e., elli…

Cited by 5SourceScholar
2023

UNTAG: Learning Generic Features for Unsupervised Type-Agnostic Deepfake Detection

ICASSP 2023accepted

This paper introduces a novel framework for unsupervised type-agnostic deepfake detection called UNTAG. Existing methods are generally trained in a supervised manner at the classification level, focusing on detecting at most two types of forgeries; thus, limiting their generalization capability acro…

Cited by 0SourceScholar
2022

Leveraging Equivariant Features for Absolute Pose Regression

CVPR 2022poster

While end-to-end approaches have achieved state-of-the-art performance in many perception tasks, they are not yet able to compete with 3D geometry-based methods in pose estimation. Moreover, absolute pose regression has been shown to be more related to image retrieval. As a result, we hypothesize th…

Cited by 35PDFScholar
2020

3d Deformation Signature for Dynamic Face Recognition

ICASSP 2020accepted

This work proposes a novel 3D Deformation Signature (3DS) to represent a 3D deformation signal for 3D Dynamic Face Recognition. 3DS is computed given a non-linear 6D-space representation which guarantees physically plausible 3D deformations. A unique deformation indicator is computed per triangle in…

Cited by 0SourceScholar
2019

Learning to Fuse Latent Representations for Multimodal Data

ICASSP 2019accepted

Multimodal learning leverages data from different modalities to improve the performance of a trained model. Typically, latent representations extracted from multimodal data are provided via direct feature fusion for end-to-end training of a deep neural network towards a specific task. However, the i…

Cited by 0SourceScholar
2019

View-invariant Action Recognition from RGB Data via 3D Pose Estimation

ICASSP 2019accepted

In this paper, we propose a novel view-invariant action recognition method using a single monocular RGB camera. View-invariance remains a very challenging topic in 2D action recognition due to the lack of 3D information in RGB images. Most successful approaches make use of the concept of knowledge t…

Cited by 0SourceScholar
2018

A Revisit of Action Detection Using Improved Trajectories

ICASSP 2018accepted

In this paper, we revisit trajectory-based action detection in a potent and non-uniform way. Improved trajectories have been proven to be an effective model for motion description in action recognition. In temporal action localization, however, this approach is not efficiently exploited. Trajectory…

Cited by 0SourceScholar
2018

Improving the Capacity of Very Deep Networks with Maxout Units

ICASSP 2018accepted

Deep neural networks inherently have large representational power for approximating complex target functions. However, models based on rectified linear units can suffer reduction in representation capacity due to dead units. Moreover, approximating very deep networks trained with dropout at test tim…

Cited by 0SourceScholar
2017

Unsupervised Vanishing Point Detection and Camera Calibration From a Single Manhattan Image With Radial Distortion

CVPR 2017poster

The article concerns the automatic calibration of a camera with radial distortion from a single image. It is known that, under the mild assumption of square pixels and zero skew, lines in the scene project into circles in the image, and three lines suffice to calibrate the camera up to an ambiguity…

Cited by 49PDFScholar