PANACEA Environment Italian monolingual corpus

Instance of: Resource Info
Description The PANACEA Environment Italian monolingual corpus was acquired in the framework of the PANACEA project (Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies), under the European Commission's Seventh Framework Programme. This corpus contains documents that were acquired from the web, were automatically detected to be in the Italian language and were automatically classified as relevant to the “Environment” domain. It was constructed in the summer of 2011 using the Focused Monolingual Crawler (FMC) developed in the context of PANACEA. The corpus contains 40,044,852 tokens, excluding tokens in a) short (smaller than 10 tokens) paragraphs, b) paragraphs that were automatically classified as boilerplate and c) paragraphs that were automatically detected in a language other than Italian. They are divided into a total of 16,159 documents that were crawled from 1,211 web sites. The dataset consists of the original crawled HTML files and the corresponding CesDoc XML files with basic metadata.
Le corpus monolingue italien PANACEA - domaine de l’environnement a été acquis dans le cadre du projet PANACEA (Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies), du septième programme-cadre de la Commission européenne. Ce corpus comprend des documents acquis depuis le web, qui ont été détectés automatiquement comme étant de langue italienne et classés automatiquement comme appartenant au domaine de l’environnement. Il a été construit en été 2011 au moyen du FMC (Focused Monolingual Crawler), développé dans le contexte de PANACEA. Le corpus contient 40 044 852 tokens, en excluant les tokens se trouvant a) dans les paragraphes courts (moins de 10 tokens), b) dans les paragraphes qui ont été classés automatiquement comme « boilerplate » et c) dans les paragraphes qui ont été détectés automatiquement comme étant dans une autre langue que l’italien. Les tokens proviennent d’un total de 16 159 documents qui ont été crawlés depuis 1 211 sites web. L’ensemble de données comprend les fichiers HTML d’origine crawlés et les fichiers XML CesDoc correspondants avec les métadonnées de base.
Language ita
Language Italian
Rights ELRA_END_USER
See Also http://metashare.elda.org/repository/browse/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697/
Source META-SHARE
Title PANACEA Environment Italian monolingual corpus
Corpus monolingue italien PANACEA - domaine de l’environnement
Type Dataset
Type Corpus
Is Is Replaced By of PANACEA Environment Italian monolingual corpus

Contact Point

Communication Info
Address 55-57 rue Brillat-Savarin
City Paris
Country France
Distribution
Access URL http://www.elda.org
Type Distribution
URL
Email [email protected]
Fax Number +1 43 14 33 30
Telephone Number +1 43 13 33 33
Type Communication Info
Zip Code 75013
Given Name Mapelli
Surname Valérie
Type Contact Person
Person
Person Info Type

Corpus Info

Corpus Text Info
Creation Info
Creation Mode Automatic
Type Creation Info
Language Info
Language Italian
Language ita
Language Name Italian
Type Language Info
Linguality Info
Linguality Type Monolingual
Type Linguality Info
Media Type Text
Size Info
Size no size available
Size Unit Other
Type Size Info Type
Text Format Info
Mime Type Plain text
Type Text Format Info
Type Corpus Text Info
Resource Type Corpus
Type Corpus Info

Distribution Info

Availability Available-restricted Use
Availability Start Date 2013-01-30 Date
License
Membership Info
Member true Boolean
Membership Institution ELRA
Type Membership Info
Permission
Action http://creativecommons.org/ns/Distribution
http://creativecommons.org/ns/CommercialUse
Constraint Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Operator Eq
Purpose Academic Use
Type Prohibition
Constraint
Permission
Restrictions Of Use
Prohibition Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Academic
Membership Info Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#membership Info2
Permission Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Prohibition Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Commercial
Membership Info
Member false Boolean
Membership Institution ELRA
Type Membership Info
Permission Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Prohibition Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Academic
Membership Info Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#membership Info
Permission Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Prohibition Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Commercial
Type Distribution
Distribution Info

Identification Info

Description The PANACEA Environment Italian monolingual corpus was acquired in the framework of the PANACEA project (Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies), under the European Commission's Seventh Framework Programme. This corpus contains documents that were acquired from the web, were automatically detected to be in the Italian language and were automatically classified as relevant to the “Environment” domain. It was constructed in the summer of 2011 using the Focused Monolingual Crawler (FMC) developed in the context of PANACEA. The corpus contains 40,044,852 tokens, excluding tokens in a) short (smaller than 10 tokens) paragraphs, b) paragraphs that were automatically classified as boilerplate and c) paragraphs that were automatically detected in a language other than Italian. They are divided into a total of 16,159 documents that were crawled from 1,211 web sites. The dataset consists of the original crawled HTML files and the corresponding CesDoc XML files with basic metadata.
Le corpus monolingue italien PANACEA - domaine de l’environnement a été acquis dans le cadre du projet PANACEA (Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies), du septième programme-cadre de la Commission européenne. Ce corpus comprend des documents acquis depuis le web, qui ont été détectés automatiquement comme étant de langue italienne et classés automatiquement comme appartenant au domaine de l’environnement. Il a été construit en été 2011 au moyen du FMC (Focused Monolingual Crawler), développé dans le contexte de PANACEA. Le corpus contient 40 044 852 tokens, en excluant les tokens se trouvant a) dans les paragraphes courts (moins de 10 tokens), b) dans les paragraphes qui ont été classés automatiquement comme « boilerplate » et c) dans les paragraphes qui ont été détectés automatiquement comme étant dans une autre langue que l’italien. Les tokens proviennent d’un total de 16 159 documents qui ont été crawlés depuis 1 211 sites web. L’ensemble de données comprend les fichiers HTML d’origine crawlés et les fichiers XML CesDoc correspondants avec les métadonnées de base.
Distribution
Access URL http://catalog.elra.info/product_info.php?products_id=1190
Type Distribution
URL
Identifier ELRA-W0069
Meta Share Id NOT_DEFINED_FOR_V2
Title PANACEA Environment Italian monolingual corpus
Corpus monolingue italien PANACEA - domaine de l’environnement
Type Identification Info

Resource Creation Info

Creation End Date 2011-01-01 Date
Funding Project
Funding Type Other
Project Name PANACEA (Platform for Automatic, Normalized Annotation and Cost-Effective Acquisition of Language Resources for Human Language Technologies)
Type Project Info Type
Type Resource Creation Info

Version Info

Has Version 1.0
Modified 2013-01-30 Date
Type Version Info

Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#metadata Info

Instance of: Catalog Record
Created 2005-05-12 Date
Primary Topic PANACEA Environment Italian monolingual corpus
Type Metadata Info

Metashare/b2b0de74de6811e2b1e400259011f6ea592d69c451df4dd1a4d960c643b80697#Header

Instance of: Catalog Record
Issued 2014-09-23T00:16:52Z Date
Primary Topic PANACEA Environment Italian monolingual corpus
Set Spec corpus:text
corpus