En aquesta pàgina trobareu tots els conjunts de dades que Softcatalà ha creat com ara corpus i diccionaris. Les dades són clau en els sistemes de lingüística computacional, i imprescindibles per a l’aprenentatge automàtic. Obrim aquests dades amb l'esperit que serveixen a tothom per crear nous projectes.
Wikimedia Commons Audio — Catalan nou
This is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or…
catalan-dictionary
En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals.
Catalan YouTube Speech Corpus
Catalan YouTube Speech Corpus This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source…
ca-text-corpus
Dataset Card for ca-text-corpus Dataset Summary Public domain corpus of Catalan text. Supported Tasks and Leaderboards This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for…
Catalan-English and Catalan-German aligned corpora to train NMT systems.
Dataset Card for Tilde-MODEL-Catalan Dataset Summary This dataset contains two dataset pairs corresponding to the Europarl corpus. Both the English and the German version are aligned with the Catalan translation, which has been obtained using Apertium's RBMT system from the…
Softcatalà website content.
Dataset Card for Softcatala-Web-Texts-Dataset Dataset Summary This repository contains Softcatala website content (articles and programs descriptions). Dataset size: articles.json contains 623 articles with 373233 words. programes.json contains 330 program descriptions with 49868 words. The license of the data is Attribution-ShareAlike…
Optimot Linguistic Data
Optimot Linguistic Data This dataset contains 4,011 entries extracted from the public Optimot linguistic consultation service of the Departament de Política Lingüística, Generalitat de Catalunya. Each record addresses a Catalan language question or linguistic topic and includes an explanation, source…
open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus Dataset Summary Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural…
Catalan-German aligned corpora to train NMT systems.
Dataset Card for Tilde-MODEL-Catalan Dataset Summary This dataset contains the German version of the Tilde-MODEL corpus aligned with a Catalan translation. The catalan text has been obtained using Apertium's RBMT system from the Spanish version. It contains 3.4M segments. Supported…