Europarl QTLeap WSD/NED corpus

Europarl QTLeap WSD/NED corpus

This corpora is part of Deliverable 5.5 of the European Commission project QTLeap FP7-ICT-2013.4.1-610516 (

The texts are sentences from the Europarl parallel corpus (Koehn, 2005). We selected the monolingual sentences from parallel corpora for the following pairs: Bulgarian-English, Czech-English, Portuguese-English and Spanish-English. The English corpus is comprised by the English side of the Spanish-English corpus.

Basque is not in Europarl. In addition, it contains the Basque and English sides of the GNOME corpus.

The texts have been automatically annotated with NLP tools, including Word Sense Disambiguation, Named Entity Disambiguation and Coreference resolution. Please check deliverable D5.6 in for more information.

Institutions involved in the annotation:
University of the Basque Country (UPV/EHU)
Faculty of Science, Univeristy of Lisbon (FCUL)
Charles University in Prague (CUNI)
Bulgarian Academy of Sciences (IICT-BAS)

You don’t have the permission to edit this resource.