Intelligent Assessment of 95598 Speech Transcription Text Quality Based on Topic Model

The quality of speech transcripts is of great significance to power system management and is an important basis for supporting subsequent data analysis. In this paper, based on the characteristics of speech transcription texts of customer service, this paper proposes an analysis method combining man...

Full description

Saved in:
Bibliographic Details
Published in:IOP conference series. Materials Science and Engineering 2019-07, Vol.563 (4), p.42001
Main Authors: Song, Bochuan, Wu, Peng, Zhang, Qiang, Chai, Bo, Gao, Yuanbo, Yang, He
Format: Article
Language:English
Subjects:
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:The quality of speech transcripts is of great significance to power system management and is an important basis for supporting subsequent data analysis. In this paper, based on the characteristics of speech transcription texts of customer service, this paper proposes an analysis method combining manual processing and latent Dirichlet allocation topic model, analyzing the transcribed texts. First, data preprocessing is performed on the State Grid's work order data, and then the text topic distribution calculation is performed by the LDA topic model, and the topic parameter is set to a total of 100 topics. Next, the unsupervised clustering of the documents is performed by the k-means method, and the similarity between the files is obtained. Finally, the quality of the data is analyzed by combining manual labeling and manual evaluation. For the first time, this paper marks and identifies the State Grid's work order analysis data, which is a pioneering work for natural language processing technology in the field of power grid.
ISSN:1757-8981
1757-899X
DOI:10.1088/1757-899X/563/4/042001