Using Generative-AI Speech-to-Text Output to Provide Automated Monitoring of Television Subtitles
This paper describes a proof-of-concept approach to monitoring timing errors and word loss in TV subtitles. It reviews previous attempts at subtitle monitoring and the problems caused to viewers by subtitle timing errors and word loss. It then introduces the use of speech to text technology and the conventions in subtitling where repetition, non-speech content and errors can make the task of aligning the speech-to-text transcript to subtitles more challenging. The paper describes the approach taken to remove non-speech content from the subtitles and transcript, along with the complex and iterative natural language processing techniques devised to ensure a sufficiently accurate alignment between the two. It then gives examples of the ways in which the results are displayed and some sample results showing the scale of problems with subtitle quality, along with examples of 24-hour plots, which are a convenient way to identify issues across a large sample of data. The paper concludes by reviewing the limits of this approach in terms of accuracy and points out the need for human oversight.
- Print ISSN
- 1545-0279
- Electronic ISSN
- 2160-2492
- Published
- 2026-07
- Content type
- Original Research
- Keywords
- accessibility, quality, monitoring, subtitles, captions
- DOI
- 10.5594/JMI.2026/TZBT7663