Two-Fold Patch Perturbation for Efficient Self-Supervised Learning in 3D Medical Imaging
Tirthajit Baruah, Kabir Jamadar, Punit Rathore
Abstract
Self-supervised pre-training has become a key paradigm for reducing annotation costs in 3D medical imaging, yet many recent approaches rely on complex objectives or incur substantial computational overhead. We propose a simple and efficient self-supervised pre-training framework for 3D medical images based on a two-fold patch-wise perturbation strategy. The method applies Bernoulli patch masking and discrete rotations, and trains a shared encoder with a three-head objective for reconstruction, perturbation localization, and rotation prediction. This design encourages spatially aware and transferable representations while remaining computationally lightweight. Experiments across diverse segmentation and classification benchmarks, including modality-shift scenarios, demonstrate consistent improvements over general self-supervised baselines and competitive or superior performance compared to recent medical self-supervised methods, while requiring substantially less memory, computation, and training time than the state-of-the-art pre-training pipelines.
BibTeX
@inproceedings{ijcai2026_twofoldpatchpert,
title = {Two-Fold Patch Perturbation for Efficient Self-Supervised Learning in 3D Medical Imaging},
author = {Tirthajit Baruah and Kabir Jamadar and Punit Rathore},
booktitle = {IJCAI 2026},
year = {2026}
}