H^2Net: Homo- and Heterogeneous Networks for Unified Segmentation
Jinyu Han, Changguang Wu, Fuming Sun, Mengyin Wang, Jinhui Tang
Abstract
Unified segmentation aims to consolidate multiple vision tasks into a single model, yet faces two core challenges: learning robust homogeneous features (e.g., shared low- and mid-level cues) to enable cross-domain knowledge transfer, while disentangling heterogeneous features (e.g., task-specific semantic objectives) to avoid negative transfer and preserve task independence. To address these challenges, we propose the Homo- and Heterogeneous Network (H2Net), a unified framework that jointly models shared homogeneous representations and task-specific heterogeneous features. Specifically, H2Net incorporates a Cross-Modal Structure Enhancement Module (CSEM), which integrates auxiliary depth priors via joint frequency–spatial cross-modal attention to strengthen task-agnostic structural representations. In addition, a Task Adapter Pool (TAP) is introduced to model task-specific heterogeneous features by assigning dedicated adapters to individual tasks, enabling task-aware feature modulation and semantic disentanglement within a shared backbone. Extensive experiments on benchmarks spanning eight tasks demonstrate the effectiveness of the proposed approach and its superior performance. Code and results will be available at \url{https://h2net-ijcai26.github.io}.
BibTeX
@inproceedings{ijcai2026_h2nethomoandhete,
title = {H^2Net: Homo- and Heterogeneous Networks for Unified Segmentation},
author = {Jinyu Han and Changguang Wu and Fuming Sun and Mengyin Wang and Jinhui Tang},
booktitle = {IJCAI 2026},
year = {2026}
}