PTPARL Corpus

9 Last view: 2024-07-09

View resource name in all available languages

Corpus PTPARL

http://catalog.elra.info/product_info.php?products_id=1179

ID:

ELRA-W0060

The PTPARL Corpus contains 1,076 texts consisting of adapted transcriptions of the Portuguese Parliament sessions. The corpus contains 1,000,441 tokens.

The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The PTPARL Corpus is a subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120.
The Corpus is delivered with the annotation manual of the CRPC corpus, a metadata file and a narrative description of the resource.

View resource description in all available languages

Le corpus PTPARL est constitué de 1 076 textes consistant en des transcriptions adaptées des sessions du Parlement portugais. Le corpus contient 1 000 441 tokens.

Le corpus est livré en un seul fichier, dans deux formats différents. La version txt contient une phrase par ligne, un numéro d’identification pour chaque texte et aucune autre annotation. Les fichier cqpweb contient un token par ligne, suivi par une étiquette de partie du discours et le lemme correspondant, avec l’annotation des chunks NP. Le corpus PTPARL est un sous-ensemble du Corpus de référence du portugais contemporain et suit le même schéma d’annotation. Pour plus d’informations sur sa préparation et son annotation, voir Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120.
Le corpus est fourni avec le manuel d’annotation du corpus de référence, un fichier de meta-données et une description narrative de la ressource.

You don’t have the permission to edit this resource.

DistributionAvailability

Available - Restricted Use

Start date: 12/05/2012

Licence

ELRA VAR

Restrictions: Commercial Use

For Members of ELRA

User Nature: Academic

ELRA END USER

Restrictions: Academic - Non Commercial Use

For Non Members of ELRA

User Nature: Academic

ELRA VAR

Restrictions: Commercial Use

For Non Members of ELRA

User Nature: Academic

ELRA END USER

Restrictions: Academic - Non Commercial Use

For Non Members of ELRA

User Nature: Commercial

ELRA VAR

Restrictions: Commercial Use

For Non Members of ELRA

User Nature: Commercial

ELRA END USER

Restrictions: Academic - Non Commercial Use

For Members of ELRA

User Nature: Academic

ELRA END USER

Restrictions: Academic - Non Commercial Use

For Members of ELRA

User Nature: Commercial

ELRA VAR

Restrictions: Commercial Use

For Members of ELRA

User Nature: Commercial

Contact Person

Mapelli Valérie

text

Monolingual text corpusLanguages

Portuguese

Linguality

Linguality type: Monolingual

Size

no size available

Metadata

Created: 05/12/2005

Version

Version: 1.0

Last Updated: 12/05/2012

People who looked at this resource also viewed the following: