Speech Activity Detection in Online Broadcast Transcription Using Deep Neural Networks and Weighted Finite State Transducers

Matějů Lukáš

Speech Activity Detection in Online Broadcast Transcription Using Deep Neural Networks and Weighted Finite State Transducers

Files

SPEECH ACTIVITY.pdf(280.46 KB)

Date

2017-01-01

Authors

Publisher

Institute of Electrical and Electronics Engineers Inc.

Abstract

In this paper, a new approach to online Speech Activity Detection (SAD) is proposed. This approach is designed for the use in a system that carries out 24/7 transcription of radio/TV broadcasts containing a large amount of non-speech segments, such as advertisements or music. To improve the robustness of detection, we adopt Deep Neural Networks (DNNs) trained on artificially created mixtures of speech and non-speech signals at desired levels of signal-to-noise ratio (SNR). An integral part of our approach is an online decoder based on Weighted Finite State Transducers (WFSTs); this decoder smooths the output from DNN. The employed transduction model is context-based, i.e., both speech and non-speech events are modeled using sequences of states. The presented experimental results show that our approach yields state-of-the-art results on standardized QUT-NOISE-TIMIT data set for SAD and, at the same time, it is capable of a) operating with low latency and b) reducing the computational demands and error rate of the target transcription system.

Subject(s)

deep neural networks, speech activity detection, weighted finite state transducers, speech recognition

Item identifier

https://dspace.tul.cz/handle/15240/31351
https://ieeexplore.ieee.org/document/7953200
https://doi.org/10.1109/ICASSP.2017.7953200

ISSN

1520-6149

ISBN

978-1-5090-4117-6

Show full item record