Corpus écrit NEMLAR

Instance of: Resource Info
Description Ce corpus a été produit dans le cadre du projet NEMLAR (http://www.nemlar.org). Deux autres ressources, produites dans le cadre du même projet, sont également disponibles : le corpus oral d’actualités radiophoniques NEMLAR (ELRA-S0219) et le corpus de synthèse de parole NEMLAR (ELRA-S0220). Le corpus écrit NEMLAR est constitué de 500 000 mots de texte arabe regroupés en 13 catégories différentes, visant à obtenir un corpus bien équilibré qui offre une représentation de la variété de traits syntaxiques, sémantiques et pragmatiques de la langue arabe moderne. Les différentes catégories sont : • Actualités politiques : 48 000 mots • Débat politique : 30 000 mots • Texte Islamique (prières et autres) : 29 000 mots • Expressions de mots communs : 8 500 mots • Textes extraits d’émissions radiophoniques : 5 500 mots • Affaires : 20 000 mots • Littérature arabe : 30 000 mots • Actualités générales : 100 000 mots • Interviews : 56 000 mots • Presse scientifique : 50 000 mots • Presse sportive : 50 000 mots • Explications d’entrées de dictionnaire : 52 000 mots • Texte du domaine juridique : 21 000 mots La période de temps des données se situe entre la fin des années 1990 jusqu’à 2005. Le corpus est fourni sous la forme de 4 versions différentes: • Texte brut • Texte entièrement voyellé • Texte comprenant une analyse lexicale de l’arabe • Texte comprenant des étiquettes pour la partie du discours Les diacritiques, l’analyse lexicale et les étiquettes pour la partie du discours ont été générées par l’outil Fassieh© de RDI. La précision de l’analyse automatique est d’environ 95%. Afin d’obtenir près de 99% de taux de précision, les linguistes ont utilisé le mode de révision visuelle de Fassieh© où le linguiste doit soit approuver la première analyse comme la plus probable (la plupart du temps) ou sélectionner une autre manuellement (pour une minorité de 4% des cas). La base de données est distribuée sur 1 CD-ROM ISO 9660. Elle a été validée par un partenaire externe et un rapport de validation est fourni.
This corpus was produced within the NEMLAR project (http://www.nemlar.org). Two other resources, produced within the same project, are also available: NEMLAR Broadcast News Speech Corpus (ELRA-S0219) and the NEMLAR Speech Synthesis Corpus (ELRA-S0220). The NEMLAR Written Corpus consists of about 500,000 words of Arabic text from 13 different categories, aiming to achieve a well-balanced corpus that offers a representation of the variety in syntactic, semantic and pragmatic features of modern Arabic language. The different categories are: • Political news: 48,000 words • Political debate: 30,000 words • Islamic text (Preaching and others): 29,000 words • Phrases of common words: 8,500 words • Text from broadcast news: 5,500 words • Business: 20,000 words • Arabic literature: 30,000 words • General news: 100,000 words • Interviews: 56,000 words • Scientific press: 50,000 words • Sports press: 50,000 words • Dictionary entries explanation: 52,000 words • Legal domain text: 21,000 words The time span of the data included goes from late 1990’s to 2005. The corpus is provided in 4 different versions: • Raw text • Fully vowelized text • Text with Arabic lexical analysis • Text with Arabic POS-tags Diacritics, lexical analysis and POS-tags were generated by RDI’s tool Fassieh©. The accuracy of the automatic analysis is around 95%. To reach about the 99% accuracy rate as defined for this corpus, the linguists used the visual revision mode of Fassieh© where the linguist has to either approve the 1st most likely analysis (most of the time) or select another one manually (in the 4% minority of the cases). The database is distributed on 1 ISO 9660 CD-ROM volume. It has been validated by an external partner and a validation report is provided.
Language ara
Language Arabic
Rights ELRA_END_USER
ELRA_VAR
See Also http://metashare.elda.org/repository/browse/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58/
Source META-SHARE
Title Corpus écrit NEMLAR
NEMLAR Written Corpus
Type Dataset
Type Corpus
Is Is Replaced By of NEMLAR Written Corpus

Contact Point

Communication Info
Address 55-57 rue Brillat-Savarin
City Paris
Country France
Distribution
Access URL http://www.elda.org
Type Distribution
URL
Email [email protected]
Fax Number +1 43 14 33 30
Telephone Number +1 43 13 33 33
Type Communication Info
Zip Code 75013
Given Name Mapelli
Surname Valérie
Type Contact Person
Person
Person Info Type

Corpus Info

Corpus Text Info
Annotation Info
Annotation Mode Mixed
Annotation Type Other
Type Annotation Info
Creation Info
Creation Mode Mixed
Type Creation Info
Language Info
Language Arabic
Language ara
Language Name Arabic
Type Language Info
Linguality Info
Linguality Type Monolingual
Type Linguality Info
Media Type Text
Size Info
Size no size available
Size Unit Other
Type Size Info Type
Type Corpus Text Info
Resource Type Corpus
Type Corpus Info

Distribution Info

Availability Available-restricted Use
Availability Start Date 2006-08-11 Date
License
Membership Info
Member false Boolean
Membership Institution ELRA
Type Membership Info
Permission
Action http://creativecommons.org/ns/Distribution
http://creativecommons.org/ns/CommercialUse
Constraint Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Operator Eq
Purpose Academic Use
Type Prohibition
Constraint
Permission
Restrictions Of Use
Prohibition Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Commercial
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_VAR_140312.pdf
Type Licence Info
User Nature Commercial
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_VAR_140312.pdf
Type Licence Info
User Nature Academic
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Prohibition Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Academic
Membership Info
Member true Boolean
Membership Institution ELRA
Type Membership Info
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_VAR_140312.pdf
Type Licence Info
User Nature Commercial
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info2
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Prohibition Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Commercial
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info2
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_VAR_140312.pdf
Type Licence Info
User Nature Academic
Membership Info Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#membership Info2
Permission Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Prohibition Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#permission
Same As http://www.elra.info/IMG/pdf_ENDUSER_140312.pdf
Type Licence Info
User Nature Academic
Type Distribution Info
Distribution

Identification Info

Description This corpus was produced within the NEMLAR project (http://www.nemlar.org). Two other resources, produced within the same project, are also available: NEMLAR Broadcast News Speech Corpus (ELRA-S0219) and the NEMLAR Speech Synthesis Corpus (ELRA-S0220). The NEMLAR Written Corpus consists of about 500,000 words of Arabic text from 13 different categories, aiming to achieve a well-balanced corpus that offers a representation of the variety in syntactic, semantic and pragmatic features of modern Arabic language. The different categories are: • Political news: 48,000 words • Political debate: 30,000 words • Islamic text (Preaching and others): 29,000 words • Phrases of common words: 8,500 words • Text from broadcast news: 5,500 words • Business: 20,000 words • Arabic literature: 30,000 words • General news: 100,000 words • Interviews: 56,000 words • Scientific press: 50,000 words • Sports press: 50,000 words • Dictionary entries explanation: 52,000 words • Legal domain text: 21,000 words The time span of the data included goes from late 1990’s to 2005. The corpus is provided in 4 different versions: • Raw text • Fully vowelized text • Text with Arabic lexical analysis • Text with Arabic POS-tags Diacritics, lexical analysis and POS-tags were generated by RDI’s tool Fassieh©. The accuracy of the automatic analysis is around 95%. To reach about the 99% accuracy rate as defined for this corpus, the linguists used the visual revision mode of Fassieh© where the linguist has to either approve the 1st most likely analysis (most of the time) or select another one manually (in the 4% minority of the cases). The database is distributed on 1 ISO 9660 CD-ROM volume. It has been validated by an external partner and a validation report is provided.
Ce corpus a été produit dans le cadre du projet NEMLAR (http://www.nemlar.org). Deux autres ressources, produites dans le cadre du même projet, sont également disponibles : le corpus oral d’actualités radiophoniques NEMLAR (ELRA-S0219) et le corpus de synthèse de parole NEMLAR (ELRA-S0220). Le corpus écrit NEMLAR est constitué de 500 000 mots de texte arabe regroupés en 13 catégories différentes, visant à obtenir un corpus bien équilibré qui offre une représentation de la variété de traits syntaxiques, sémantiques et pragmatiques de la langue arabe moderne. Les différentes catégories sont : • Actualités politiques : 48 000 mots • Débat politique : 30 000 mots • Texte Islamique (prières et autres) : 29 000 mots • Expressions de mots communs : 8 500 mots • Textes extraits d’émissions radiophoniques : 5 500 mots • Affaires : 20 000 mots • Littérature arabe : 30 000 mots • Actualités générales : 100 000 mots • Interviews : 56 000 mots • Presse scientifique : 50 000 mots • Presse sportive : 50 000 mots • Explications d’entrées de dictionnaire : 52 000 mots • Texte du domaine juridique : 21 000 mots La période de temps des données se situe entre la fin des années 1990 jusqu’à 2005. Le corpus est fourni sous la forme de 4 versions différentes: • Texte brut • Texte entièrement voyellé • Texte comprenant une analyse lexicale de l’arabe • Texte comprenant des étiquettes pour la partie du discours Les diacritiques, l’analyse lexicale et les étiquettes pour la partie du discours ont été générées par l’outil Fassieh© de RDI. La précision de l’analyse automatique est d’environ 95%. Afin d’obtenir près de 99% de taux de précision, les linguistes ont utilisé le mode de révision visuelle de Fassieh© où le linguiste doit soit approuver la première analyse comme la plus probable (la plupart du temps) ou sélectionner une autre manuellement (pour une minorité de 4% des cas). La base de données est distribuée sur 1 CD-ROM ISO 9660. Elle a été validée par un partenaire externe et un rapport de validation est fourni.
Distribution
Access URL http://catalog.elra.info/product_info.php?products_id=873
Type Distribution
URL
Identifier ELRA-W0042
Meta Share Id NOT_DEFINED_FOR_V2
Title NEMLAR Written Corpus
Corpus écrit NEMLAR
Type Identification Info

Resource Creation Info

Funding Project
Funding Type Eu Funds
Project Name NEMLAR (Network for Euro-Mediterranean LAnguage Resources)
Type Project Info Type
Type Resource Creation Info

Version Info

Has Version 1.0
Modified 2007-02-22 Date
Type Version Info

Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#Header

Instance of: Catalog Record
Issued 2014-09-23T00:23:03Z Date
Primary Topic NEMLAR Written Corpus
Set Spec corpus:text
corpus

Metashare/baf9e9a0de6711e2b1e400259011f6eaa11bfe7fefd041b9b094afac6a99ab58#metadata Info

Instance of: Catalog Record
Created 2005-05-12 Date
Primary Topic NEMLAR Written Corpus
Type Metadata Info