Catalan YouTube Speech Corpus

Compartiu

Descripció

Catalan YouTube Speech Corpus

This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source video's reuse license. It was built and published by Softcatalà, the volunteer organization behind free/open-source Catalan-language software.

Homepage: https://www.softcatala.org/… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-youtube-speech.

Adreça de descàrrega:

https://huggingface.co/datasets/softcatala/catalan-youtube-speech
Autor:

Softcatalà

Hugging Face:

598 baixades (darrers 30 dies)

3 «m'agrada»

Llicència:

MIT License