Efficient and Expandable Token-Level Approach for Multi-Domain Sensitive Information Classification
Hongyi Li, Jiawei Ye, Jie Wu, Lijun Zu
Abstract
Incorporating privacy regulations and business requirements, enterprises should securely manage unstructured textual data from diverse domains. Sensitive information classification is a critical component of data security, but it poses challenges due to complex textual contexts. With the ever-increasing data volumes, rule-based and lexicon-based classification methods require substantial maintenance and become inefficient. Additionally, supervised named entity recognition is not feasible for multi-domain sensitive information classification due to numerous type-specific annotations and the scarcity of training data. This work presents ToSIC, a new sensitive information classification approach to tackle these problems. Specifically, ToSIC simplifies multi-domain type labeling through a hierarchical structure and employs a rapid learning mechanism based on prototype distance classification to mitigate extensive training data. Experiments indicate that ToSIC outperforms existing methods by significantly reducing inference time and error rates while ensuring classification performance. Moreover, ToSIC demonstrates exceptional expandability to unknown types, showcasing potential value in practical applications.
BibTeX
@inproceedings{icassp2025_efficientandexpa,
title = {Efficient and Expandable Token-Level Approach for Multi-Domain Sensitive Information Classification},
author = {Hongyi Li and Jiawei Ye and Jie Wu and Lijun Zu},
booktitle = {ICASSP 2025},
year = {2025}
}