Descripció
Dataset Card for ca-text-corpus
Dataset Summary
Public domain corpus of Catalan text.
Supported Tasks and Leaderboards
This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard.
Languages
Catalan (ca).
Dataset Structure
Data Instances
Each… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/ca_text_corpus.
Adreça de descàrrega:
https://huggingface.co/datasets/softcatala/ca_text_corpus| property | value | ||||||
|---|---|---|---|---|---|---|---|
| name | ca-text-corpus |
||||||
| description | |
||||||
| license |
|
||||||
| sameAs | https://www.softcatala.org/dades-obertes/ca_text_corpus/ |
||||||
| url | https://huggingface.co/datasets/softcatala/ca_text_corpus |