Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis
Peixin Chen, Wu Guo, Lirong Dai, Zhenhua Ling
Abstract
In recent years, neural networks (NN) have achieved remarkable performance improvement in text classification due to their powerful ability to encode discriminative features by incorporating label information into model training. Inspired by the success of NN in text classification, we propose a pseudo-supervised neural network approach for text clustering. The neural network is trained in a supervised fashion with pseudo-labels, which are provided by the cluster labels of pre-clustering on unsupervised document representations. To enhance the quality of pseudo-labels, a consensus analysis is employed to select training samples for the neural network. The experimental results demonstrate that the proposed approach can improve the clustering performance significantly.
BibTeX
@inproceedings{icassp2018_pseudosupervised,
title = {Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis},
author = {Peixin Chen and Wu Guo and Lirong Dai and Zhenhua Ling},
booktitle = {ICASSP 2018},
year = {2018}
}