← Search

Ajmal Mian

46 accepted papers

2026

Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors

CVPR 2026

The performance of neural network models deteriorates due to their unreliable behavior on non-robust features of corrupted samples. Owing to their opaque nature, rectifying models to address this problem often necessitates arduous data cleaning and model retraining, resulting in huge computational a

Cited by 0SourceScholar
2026

Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D Generation

CVPR 2026

3D shape generation has become increasingly important for graphics and vision applications. Current part-aware 3D generation usually overlooks hierarchical part relations or inefficiently encodes multi-level semantics in Euclidean space. Thus we propose a novel framework for hierarchical and efficie

Cited by 0SourceScholar
2026

QPoint: End-to-End Lightweight Point Cloud Processing via Robust Quaternion Feature Learning

ICML 2026poster

The inherent sparsity, lack of structure, and rotation sensitivity of point clouds often lead to high computational and parameter cost in robust feature learning. To address these problems, we present QPoint, a lightweight framework that leverages robust quaternion feature learning. QPoint incorpora…

Cited by 0SourceScholar
2026

RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape Retrieval

CVPR 2026

3D assets have rapidly expanded in quantity and diversity due to the growing popularity of virtual reality and gaming. As a result, text-to-shape retrieval has become essential in facilitating intuitive search within large repositories. However, existing methods require canonical poses and support f

Cited by 0SourcecodeScholar
2026

SpikeNet: Sparse Spike-Driven Mask Vector Transformer for Energy-Efficient and Stable Spiking Point Cloud Processing

ICML 2026poster

The unordered nature of point cloud data poses significant challenges to conventional architectures primarily designed for structured data. Spiking neural networks (SNN), by virtue of their inherent sparsity and dynamics, are particularly well-suited for processing point clouds to effectively extrac…

Cited by 0SourceScholar
2025

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

IJCAI 2025

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieve

Cited by 0SourcePDFScholar
2025

Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation

CVPR 2025poster

3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework des…

Cited by 0SourcePDFScholar
2025

LanCOPE: Language-Guided Category-Level Object Pose Estimation From a Single RGB Image

RA-L 2025

Monocular RGB-based category-level object pose estimation is more practical and cost-effective for robotics. However, existing methods do not fully exploit the rich semantic and contextual information in multimodal data (e.g. language) that provides additional object attributes to guide the model in

Cited by 1SourceScholar
2025

MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion Model

ICRA 2025

Object pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects.

Cited by 5SourcecodeScholar
2025

Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes

IROS 2025

Fusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection

Cited by 3SourcecodeScholar
2025

Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition

CVPR 2025poster

Recent open-world representation learning approaches have leveraged CLIP to enable zero-shot 3D object recognition. However, performance on real point clouds with occlusions still falls short due to unrealistic pretraining settings. Additionally, these methods incur high inference costs because they…

2025

RGB-Based Category-Level Object Pose Estimation via Depth Recovery and Adaptive Refinement

RA-L 2025

Category-level pose estimation methods have received widespread attention as they can be generalized to intra-class unseen objects. Although RGB-D-based category-level methods have made significant progress, reliance on depth image limits practical application. RGB-based methods offer a more practic

Cited by 3SourceScholar
2024

Efficient Hyperparameter Optimization with Adaptive Fidelity Identification

CVPR 2024poster

Hyperparameter Optimization and Neural Architecture Search are powerful in attaining state-of-the-art machine learning models with Bayesian Optimization (BO) standing out as a mainstream method. Extending BO into the multi-fidelity setting has been an emerging research topic in this field but faces…

2024

External Knowledge Enhanced 3D Scene Generation from Sketch

ECCV 2024poster

"Generating realistic 3D scenes is challenging due to the complexity of room layouts and object geometries. We propose a sketch based knowledge enhanced diffusion architecture (SEK) for generating customized, diverse, and plausible 3D scenes. SEK conditions the denoising process with a hand-drawn sk…

Cited by 6SourcePDFScholar
2024

L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream

CVPR 2024poster

3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper we aim to handle a real-time multi-task for 6-DoF pose tracking of u…

Cited by 0SourcePDFScholar
2024

Regulating Model Reliance on Non-Robust Features by Smoothing Input Marginal Density

ECCV 2024poster

"Trustworthy machine learning necessitates meticulous regulation of model reliance on non-robust features. We propose a framework to delineate and regulate such features by attributing model predictions to the input. Within our approach, robust feature attributions exhibit a certain consistency, whi…

2024

Sparse Points to Dense Clouds: Enhancing 3D Detection with Limited LiDAR Data

IROS 2024

3D detection is a critical task that enables machines to identify and locate objects in three-dimensional space. It has a broad range of applications in several fields, including autonomous driving, robotics and augmented reality. Monocular 3D detection is attractive as it requires only a single cam

Cited by 6SourceScholar
2023

3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud

CVPR 2023poster

In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide…

2023

3DMODT: Attention-Guided Affinities for Joint Detection & Tracking in 3D Point Clouds

ICRA 2023poster

We propose a method for joint detection and tracking of multiple objects in 3D point clouds, a task conventionally treated as a two-step process comprising object detection followed by data association. Our method embeds both steps into a single end-to-end trainable network eliminating the dependenc…

Cited by 10SourceScholar
2023

Contrastive Self-Supervised Learning Leads to Higher Adversarial Susceptibility

AAAI 2023technical

Contrastive self-supervised learning (CSL) has managed to match or surpass the performance of supervised learning in image and video classification. However, it is still largely unknown if the nature of the representations induced by the two learning paradigms is similar. We investigate this under t…

Cited by 6SourcePDFScholar
2023

DANet: Density Adaptive Convolutional Network With Interactive Attention for 3D Point Clouds

RA-L 2023

Local features and contextual dependencies are crucial for 3D point cloud analysis. Many works have been devoted to designing better local convolutional kernels that exploit the contextual dependencies. However, current point convolutions lack robustness to varying point cloud density. Moreover, con

Cited by 7SourceScholar
2023

Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation

ICCV 2023poster

Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently amb…

Cited by 32PDFScholar
2023

Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud Maps

ICRA 2023poster

Precise localization is critical for autonomous vehicles. We present a self-supervised learning method that employs transformers for the first time for the task of outdoor localization using LiDAR data. We propose a pre-text task that reorganizes the slices of a 360° LiDAR scan to leverage its axial…

Cited by 6SourceScholar
2023

Spectrum-guided Multi-granularity Referring Video Object Segmentation

ICCV 2023poster

Current referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perc…

Cited by 54PDFcodeScholar
2023

UnLoc: A Universal Localization Method for Autonomous Vehicles using LiDAR, Radar and/or Camera Input

IROS 2023poster

Localization is a fundamental task in robotics for autonomous navigation. Existing localization methods rely on a single input data modality or train several computational models to process different modalities. This leads to stringent computational requirements and sub-optimal results that fail to…

Cited by 3SourcecodeScholar
2022

Learning From Pixel-Level Noisy Label: A New Perspective for Light Field Saliency Detection

CVPR 2022poster

Saliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obta…

Cited by 25PDFcodeScholar
2022

Self Supervised Learning for Multiple Object Tracking in 3D Point Clouds

IROS 2022poster

Multiple object tracking in 3D point clouds has applications in mobile robots and autonomous driving. This is a challenging problem due to the sparse nature of the point clouds and the added difficulty of annotation in 3D for supervised learning. To overcome these challenges, we propose a neural net…

Cited by 5SourceScholar
2022

UniCon: Combating Label Noise Through Uniform Selection and Contrastive Learning

CVPR 2022poster

Supervised deep learning methods require a large repository of annotated data; hence, label noise is inevitable. Training with such noisy data negatively impacts the generalization performance of deep neural networks. To combat label noise, recent state-of-the-art methods employ some sort of sample…

Cited by 188PDFcodeScholar
2021

CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image Saliency

CVPR 2021poster

Backpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, class-insensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, res…

Cited by 69PDFcodeScholar
2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2020

Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking

ECCV 2020poster

Deep learning based Multiple Object Tracking (MOT) currently relies on off-the-shelf detectors for tracking-by-detection. This results in deep models that are detector biased and evaluations that are detector influenced. To resolve this issue, we introduce Deep Motion Modeling Network (DMM-Net) that…

Cited by 72SourcePDFScholar
2019

Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning

CVPR 2019poster

Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through…

Cited by 304PDFScholar
2018

3D Face Reconstruction from Light Field Images: A Model-free Approach

ECCV 2018poster

Reconstructing 3D facial geometry from a single RGB image has recently instigated wide research interest. However, it is still an ill-posed problem and most methods rely on prior models hence undermining the accuracy of the recovered 3D faces. In this paper, we exploit the Epipolar Plane Images (EPI…

Cited by 50SourcePDFScholar
2015

Bayesian Sparse Representation for Hyperspectral Image Super Resolution

CVPR 2015poster

Despite the proven efficacy of hyperspectral imaging in many computer vision tasks, its widespread use is hindered by its low spatial resolution, resulting from hardware limitations. We propose a hyperspectral image super resolution approach that fuses a high resolution image with the low resolution…

Cited by 329SourcePDFScholar
2015

Shape-Based Automatic Detection of a Large Number of 3D Facial Landmarks

CVPR 2015poster

We present an algorithm for automatic detection of a large number of anthropometric landmarks on 3D faces. Our approach does not use texture and is completely shape based in order to detect landmarks that are morphologically significant. The proposed algorithm evolves level set curves with adaptive…

Cited by 80SourcePDFScholar