Speech corpus

A speech corpus (or spoken corpus) is a database of speech audio files and text transcriptions.

In speech technology, speech corpora are used, among other things, to create acoustic models (which can then be used with a speech recognition or speaker identification engine).

[1] In linguistics, spoken corpora are used to do research into phonetic, conversation analysis, dialectology and other fields.

Corpora is the plural of corpus (i.e. it is many such databases).

This article about a digital library is a stub.