← Search

Sajid Javed

8 accepted papers

2026

Efficient Joint Estimation of Optical Flow and Stereo Disparity With Event Cameras

RA-L 2026

Optical flow and stereo disparity, both are fundamental in the perception pipeline of robotic systems, enabling 3D understanding and motion estimation of the robot itself and the dynamic objects around. In this context, event cameras offer great potential to reduce latency and improve the efficiency

Cited by 0SourcecodeScholar
2026

FlickerTac: Flickering LED Driven Photometric Stereo for Event Vision-Based Tactile Sensors

RA-L 2026

Tactile sensing at high speed and resolution is critical for robotic perception and control. Existing vision-based tactile sensors (VBTSs) achieve high spatial accuracy but suffer from degraded performance in dynamic interactions due to motion blur and low frame rates, limiting their effectiveness i

Cited by 0SourceScholar
2026

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

CVPR 2026

Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language Models (MLLMs) typically compress an entire WSI into a single

Cited by 0SourcecodeScholar
2026

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

CVPR 2026

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent reference tracking. Existing video MLLMs often rely on a stati

Cited by 0SourcecodeScholar
2025

Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation

CVPR 2025poster

In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tis…

2024

CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment

CVPR 2024poster

This paper proposes Comprehensive Pathology Language Image Pre-training (CPLIP) a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision language models by leveraging extens…

2023

DFR-FastMOT: Detection Failure Resistant Tracker for Fast Multi-Object Tracking Based on Sensor Fusion

ICRA 2023poster

Persistent multi-object tracking (MOT) allows autonomous vehicles to navigate safely in highly dynamic environments. One of the well-known challenges in MOT is object occlusion when an object becomes unobservant for subsequent frames. The current MOT methods store objects information, such as trajec…

Cited by 10SourcecodeScholar
2023

Higher-Order Sparse Convolutions in Graph Neural Networks

ICASSP 2023accepted

Graph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this w…

Cited by 0SourceScholar