← Search

NAVEED AKHTAR

33 accepted papers

2026

Attribution-Guided Model Rectification of Unreliable Neural Network Behaviors

CVPR 2026

The performance of neural network models deteriorates due to their unreliable behavior on non-robust features of corrupted samples. Owing to their opaque nature, rectifying models to address this problem often necessitates arduous data cleaning and model retraining, resulting in huge computational a

Cited by 3SourceScholar
2026

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

ICLR 2026poster

Object hallucination in Multimodal Large Language Models (MLLMs) is a persistent failure mode that causes the model to perceive objects absent in the image. This weakness of MLLMs is currently studied using static benchmarks with fixed visual scenarios, which preempts the possibility of uncovering m…

Cited by 0SourcecodeScholar
2025

Beyond Human Perception: Understanding Multi-Object World from Monocular View

CVPR 2025poster

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often availab…

2025

CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation

NeurIPS 2025poster

Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, th…

Cited by 0SourceScholar
2025

DDB: Diffusion Driven Balancing to Address Spurious Correlations

ICCV 2025poster

Deep neural networks trained with Empirical Risk Minimization (ERM) perform well when both training and test data come from the same domain, but they often fail to generalize to out-of-distribution samples. In image classification, these models may rely on spurious correlations that often exist betw…

2025

GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector

CVPR 2025poster

We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is…

2025

Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking

CVPR 2025poster

Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal wi…

2025

Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes

IROS 2025

Fusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and RGB input has started gaining traction. However, effective integration of these modalities for precise object detection

Cited by 3SourcecodeScholar
2025

Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control

CVPR 2025poster

Ethical issues around text-to-image (T2I) models demand a comprehensive control over the generative content. Existing techniques addressing these issues for responsible T2I models aim for the generated content to be fair and safe (non-violent/explicit). However, these methods remain bounded to han…

2025

RID-Net: A Hybrid MLP-Transformer Network for Robust Point Cloud Registration

RA-L 2025

The robustness of correspondence-based point cloud registration relies on transformation invariance and intrinsic distinctiveness of the descriptors computed for registration. However, for challenging scenarios with different objects having similar local geometry and low point cloud overlap, existin

Cited by 0SourceScholar
2025

Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

AAAI 2025technical

3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to ac…

2024

Fine-Grained Semantic Information Preservation and Misclassification-Aware Loss for 3D Point Cloud

RA-L 2024

Encoder-Decoder structure is a popular choice in point cloud processing for dense multi-classification tasks, e.g., 3D semantic segmentation. Though existing techniques that follow this structure achieve high performance, they are known to suffer from fine-grained information loss, especially when t

Cited by 0SourceScholar
2024

Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding

AAAI 2024technical

Multi-Layer Perceptron (MLP) models are the bedrock of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength. We first develop an “abstraction and refinement” (ABS-REF) view for the neural modeling of point clouds. This view elucidates…

2024

Regulating Model Reliance on Non-Robust Features by Smoothing Input Marginal Density

ECCV 2024poster

"Trustworthy machine learning necessitates meticulous regulation of model reliance on non-robust features. We propose a framework to delineate and regulate such features by attributing model predictions to the input. Within our approach, robust feature attributions exhibit a certain consistency, whi…

2023

AShapeFormer: Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers

CVPR 2023poster

3D object detection techniques commonly follow a pipeline that aggregates predicted object central point features to compute candidate points. However, these candidate points contain only positional information, largely ignoring the object-level shape information. This eventually leads to sub-optima…

2023

Contrastive Self-Supervised Learning Leads to Higher Adversarial Susceptibility

AAAI 2023technical

Contrastive self-supervised learning (CSL) has managed to match or surpass the performance of supervised learning in image and video classification. However, it is still largely unknown if the nature of the representations induced by the two learning paradigms is similar. We investigate this under t…

Cited by 6SourcePDFScholar
2023

Re-calibrating Feature Attributions for Model Interpretation

ICLR 2023top-25%

The ability to interpret machine learning models is critical for high-stakes applications. Due to its desirable theoretical properties, path integration is a widely used scheme for feature attribution to interpret model predictions. However, the methods implementing this scheme currently rely on abs…

Cited by 13SourcePDFScholar
2023

Rethinking Interpretation: Input-Agnostic Saliency Mapping of Deep Visual Classifiers

AAAI 2023technical

Saliency methods provide post-hoc model interpretation by attributing input features to the model outputs. Current methods mainly achieve this using a single input sample, thereby failing to answer input-independent inquiries about the model. We also show that input-specific saliency mapping is intr…

Cited by 7SourcePDFScholar
2023

Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud Maps

ICRA 2023poster

Precise localization is critical for autonomous vehicles. We present a self-supervised learning method that employs transformers for the first time for the task of outdoor localization using LiDAR data. We propose a pre-text task that reorganizes the slices of a 360° LiDAR scan to leverage its axial…

Cited by 6SourceScholar
2023

UnLoc: A Universal Localization Method for Autonomous Vehicles using LiDAR, Radar and/or Camera Input

IROS 2023poster

Localization is a fundamental task in robotics for autonomous navigation. Existing localization methods rely on a single input data modality or train several computational models to process different modalities. This leads to stringent computational requirements and sub-optimal results that fail to…

Cited by 3SourcecodeScholar
2022

Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point Clouds

CVPR 2022poster

Point cloud scene flow estimation is of practical importance for dynamic scene navigation in autonomous driving. Since scene flow labels are hard to obtain, current methods train their models on synthetic data and transfer them to real scenes. However, large disparities between existing synthetic da…

Cited by 37PDFcodeScholar
2021

CAMERAS: Enhanced Resolution and Sanity Preserving Class Activation Mapping for Image Saliency

CVPR 2021poster

Backpropagation image saliency aims at explaining model predictions by estimating model-centric importance of individual pixels in the input. However, class-insensitivity of the earlier layers in a network only allows saliency computation with low resolution activation maps of the deeper layers, res…

Cited by 69PDFcodeScholar
2020

Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking

ECCV 2020poster

Deep learning based Multiple Object Tracking (MOT) currently relies on off-the-shelf detectors for tracking-by-detection. This results in deep models that are detector biased and evaluations that are detector influenced. To resolve this issue, we introduce Deep Motion Modeling Network (DMM-Net) that…

Cited by 72SourcePDFScholar
2019

Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning

CVPR 2019poster

Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through…

Cited by 304PDFScholar
2015

Bayesian Sparse Representation for Hyperspectral Image Super Resolution

CVPR 2015poster

Despite the proven efficacy of hyperspectral imaging in many computer vision tasks, its widespread use is hindered by its low spatial resolution, resulting from hardware limitations. We propose a hyperspectral image super resolution approach that fuses a high resolution image with the low resolution…

Cited by 329SourcePDFScholar