Catalan-English and Catalan-German aligned corpora to train NMT systems.

Compartiu

Descripció

Dataset Card for Tilde-MODEL-Catalan

Dataset Summary

This dataset contains two dataset pairs corresponding to the Europarl corpus. Both the English and the German version are aligned with the Catalan translation, which has been obtained using Apertium's RBMT system from the Spanish version of the Spanish-English alignment. Catalan-German alignment has been obtained using this alignment finder from de-en and ca-en.

Catalan-English: 1 965 735 segments.… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/Europarl-catalan.

Adreça de descàrrega:

https://huggingface.co/datasets/softcatala/Europarl-catalan
Autor:

Softcatalà

Hugging Face:

56 baixades (darrers 30 dies)

1 «m'agrada»