← Search

M. Saquib Sarfraz

10 accepted papers

2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

ICLR 2026poster

Despite substantial progress in video understanding, most existing datasets are limited to Earth’s gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-rob…

Cited by 0SourcecodeScholar
2025

Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

ICCV 2025poster

In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than in…

2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

NeurIPS 2024poster

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and ac…

2024

Improving Single Domain-Generalized Object Detection: A Focus on Diversification and Alignment

CVPR 2024poster

In this work we tackle the problem of domain generalization for object detection specifically focusing on the scenario where only a single source domain is available. We propose an effective approach that involves two key steps: diversifying the source domain and aligning detections based on class p…

2024

Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle Activations

NeurIPS 2024poster

Exploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of dat…

Cited by 3SourcePDFScholar
2024

Navigating Open Set Scenarios for Skeleton-Based Action Recognition

AAAI 2024technical

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and t…

2024

Position: Quo Vadis, Unsupervised Time Series Anomaly Detection?

ICML 2024poster

The current state of machine learning scholarship in Timeseries Anomaly Detection (TAD) is plagued by the persistent use of flawed evaluation metrics, inconsistent benchmarking practices, and a lack of proper justification for the choices made in novel deep learning-based model designs. Our paper pr…

2022

Towards Improving Calibration in Object Detection Under Domain Shift

NeurIPS 2022accept

With deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibrati…

Cited by 22SourcePDFScholar
2021

SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection

NeurIPS 2021poster

We study adapting trained object detectors to unseen domains manifesting significant variations of object appearance, viewpoints and backgrounds. Most current methods align domains by either using image or instance-level feature alignment in an adversarial fashion. This often suffers due to the pres…

Cited by 84SourcePDFScholar
2018

A Pose-Sensitive Embedding for Person Re-Identification With Expanded Cross Neighborhood Re-Ranking

CVPR 2018poster

Person re-identification is a challenging retrieval task that requires matching a person’s acquired image across non-overlapping camera views. In this paper we propose an effective approach that incorporates both the fine and coarse pose information of the person to learn a discrim- inative embeddin…

Cited by 604SourcePDFScholar