Dades Obertes

Compartiu

En aquesta pàgina trobareu tots els conjunts de dades que Softcatalà ha creat com ara corpus i diccionaris. Les dades són clau en els sistemes de lingüística computacional, i imprescindibles per a l’aprenentatge automàtic. Obrim aquests dades amb l'esperit que serveixen a tothom per crear nous projectes.

Wikimedia Commons Audio — Catalan nou

This is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or…

479 baixades Altres condicions Baixa les dades


catalan-dictionary

En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals.

1.135 baixades 3 «m'agrada» GNU General Public License v2.0 Baixa les dades


Catalan YouTube Speech Corpus

Catalan YouTube Speech Corpus This dataset contains 231,684 short audio clips of spontaneous Catalan speech, automatically extracted from public YouTube videos. Each clip is paired with two independent machine-generated transcription candidates, along with speaker gender, clip timing, and the source…

598 baixades 3 «m'agrada» MIT License Baixa les dades


ca-text-corpus

Dataset Card for ca-text-corpus Dataset Summary Public domain corpus of Catalan text. Supported Tasks and Leaderboards This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for…

57 baixades CC0 1.0 (domini públic) Baixa les dades


Catalan-English and Catalan-German aligned corpora to train NMT systems.

Dataset Card for Tilde-MODEL-Catalan Dataset Summary This dataset contains two dataset pairs corresponding to the Europarl corpus. Both the English and the German version are aligned with the Catalan translation, which has been obtained using Apertium's RBMT system from the…

56 baixades 1 «m'agrada» Creative Commons BY 4.0 Baixa les dades


Softcatalà website content.

Dataset Card for Softcatala-Web-Texts-Dataset Dataset Summary This repository contains Softcatala website content (articles and programs descriptions). Dataset size: articles.json contains 623 articles with 373233 words. programes.json contains 330 program descriptions with 49868 words. The license of the data is Attribution-ShareAlike…

50 baixades Creative Commons BY-SA 4.0 Baixa les dades


Optimot Linguistic Data

Optimot Linguistic Data This dataset contains 4,011 entries extracted from the public Optimot linguistic consultation service of the Departament de Política Lingüística, Generalitat de Catalunya. Each record addresses a Catalan language question or linguistic topic and includes an explanation, source…

42 baixades Creative Commons BY 4.0 Baixa les dades


open-source-english-catalan-corpus

Dataset Card for open-source-english-catalan-corpus Dataset Summary Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural…

39 baixades 1 «m'agrada» GNU General Public License v3.0 Baixa les dades


Catalan-German aligned corpora to train NMT systems.

Dataset Card for Tilde-MODEL-Catalan Dataset Summary This dataset contains the German version of the Tilde-MODEL corpus aligned with a Catalan translation. The catalan text has been obtained using Apertium's RBMT system from the Spanish version. It contains 3.4M segments. Supported…

26 baixades Creative Commons BY 4.0 Baixa les dades