← Search

Waqas Sultani

8 accepted papers

2026

GeoFlow: Real-Time Fine-Grained Cross-View Geolocalization via Iterative Flow Prediction

CVPR 2026

Accurate and fast localization is vital for safe autonomous navigation in GPS-denied areas. Fine-Grained Cross-View Geolocalization (FG-CVG) aims to estimate the precise 2-Degree-of-Freedom (2-DoF) location of a ground image relative to a satellite image. However, current methods force a difficult t

Cited by 0SourcecodeScholar
2023

Cross-View Geo-Localization via Learning Disentangled Geometric Layout Correspondence

AAAI 2023technical

Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficu…

2023

TransVisDrone: Spatio-Temporal Transformer for Vision-based Drone-to-Drone Detection in Aerial Videos

ICRA 2023poster

Drone-to-drone detection using visual feed has crucial applications, such as detecting drone collisions, detecting drone attacks, or coordinating flight with other drones. However, existing methods are computationally costly, follow non-end-to-end optimization, and have complex multi-stage pipelines…

Cited by 36SourcecodeScholar
2022

Towards Low-Cost and Efficient Malaria Detection

CVPR 2022poster

Malaria, a fatal but curable disease claims hundreds of thousands of lives every year. Early and correct diagnosis is vital to avoid health complexities, however, it depends upon the availability of costly microscopes and trained experts to analyze blood-smear slides. Deep learning-based methods hav…

Cited by 24PDFScholar
2016

What If We Do Not Have Multiple Videos of the Same Action? -- Video Action Localization Using Web Images

CVPR 2016poster

This paper tackles the problem of spatio-temporal action localization in a video without assuming the availability of multiple videos or any prior annotations. Action is localized by employing images downloaded from internet using action name. Given web images, we first mitigate image noise using r…

Cited by 41PDFScholar