SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from varia…