Bulgarian-X language Parallel Corpus

Instance of: Resource Info
Description The Bulgarian-X language Parallel Corpus (Bul-X-Cor) is a part of the Bulgarian National Corpus (BulNC). The Bulgarian National Corpus is designed as a uniform framework for texts of different modality (written - spoken), period (synchronic - diachronic), and number of languages (monolingual - parallel where one of the counterparts is Bulgarian). Any X-language in the corpus is equally treated with respect to the text type diversity and balance, metadata description scheme, preprocessing and annotation, search engine queries and data storage format. Bulgarian-X Language Parallel Corpus includes parallel corpora of 48 languages – English, German, French, Slavic and Balkan languages, as well as other European and non-European languages. The parallel corpora represent only texts which have a Bulgarian correspondence – either the original is in Bulgarian, there is a Bulgarian translation, or both texts are translations from a third language. As of January 2013, the Bulgarian-X Language Parallel Corpus contains 4.2 billion tokens, comprising the biggest parallel corpus of Bulgarian. Languages are not equally represented: the largest parallel corpus is the Bulgarian-English parallel corpus (280.8 and 283.1 million words for Bulgarian and English respectively); there are 18 other corpora of over 200 million tokens per language, 2 parallel corpora between 100 and 200 million tokens per language, 11 parallel corpora of size in the range 5-15 million tokens per language, and the rest 15 are below 1 million, with the smallest corpus being Japanese with 50,000 tokens. Each parallel subcorpus within Bul-X-Cor mirrors the structure of BulNC. The structure, data formatting and text description follow the model of BulNC. All Bulgarian texts in BulNC and English texts in Bul-X-Cor are supplied with extensive metadata description compliant with the well established standards. The Bulgarian-English parallel corpus is supplied as well with annotation on various levels while the annotation of other languages has just started. Main applications of parallel corpora are in the field of computational linguistics: machine translation, developing bilingual lexical resources (dictionaries), etc. The benefits of the parallel corpora increase if they are annotated. The Bulgarian-X Language Parallel Corpus Collocations service is a web service for collocations search and different types of statistics over the Bulgarian-X Language Parallel Corpus. The service employs the free of charge NoSketchEngine, a system for corpora processing that combines Manatee and Bonito. The Collocation service is a RESTful webservice, supporting complicated queries through http. Example: http://dcl.bas.bg/collocations/?cmd=collocations&word=нет user: bulnc pass: bulnc The query returns the collocations of a given word in the NoSketchEngine format. The system also supports additional arguments, namely all that are accepted by NoSketchEngine, provided with default values and an optional language identificator. The following example restricts the statistics to Bulgarian: http://dcl.bas.bg/collocations/?cmd=collocations&word=нет&lang=bg
Language en
he
pl
da
ky
uk
tk
es
it
bs
fi
kk
sr
el
ro
de
bg
ar
mt
lt
is
sl
eu
ga
no
sk
fr
hu
hr
ca
zh
az
mn
ja
ru
nl
sq
sv
et
cs
pt
mk
tr
ka
lv
tg
hy
Language Russian
Bulgarian
Finnish
French
Italian
Basque
Romanian
Japanese
Spanish
Greek (modern)
Portuguese
Icelandic
Danish
Slovak
Dutch
English
Polish
Swedish
German
Norwegian
Catalan
Serbian
Chinese
Czech
Rights other
See Also http://metashare.elda.org/repository/browse/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b/
Source META-SHARE
Title Bulgarian-X language Parallel Corpus
Type Dataset
Type Corpus

Contact Point

Affiliation
Communication Info
Distribution
Access URL http://dcl.bas.bg
Type Distribution
URL
Email [email protected]
Type Communication Info
Department Name Department of Computational Linguistics
Organization Name Institute for Bulgarian Language
Organization Short Name IBL
Type Organization Info Type
Communication Info Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#communication Info2
Given Name Ivelina
Position Affiliated researcher
Surname Stoyanova
Type Contact Person
Person
Person Info Type

Corpus Info

Corpus Text Info
Language Info
Language bg
en
Language Bulgarian
English
Language Name Bulgarian
English
Size Per Language
Size 260,681,821
Size Unit Tokens
Type Size Info Type
Type Language Info
Linguality Info
Linguality Type Multilingual
Multilinguality Type Parallel
Type Linguality Info
Media Type Text
Size Info
Size 1,202,209,147 Tokens
Size Unit Tokens
Type Size Info Type
Type Corpus Text Info
Annotation Info
Annotation Type Segmentation
Segmentation Level Sentence
Type Annotation Info
Annotation Type Segmentation
Segmentation Level Word
Type Annotation Info
Annotation Type Morphosyntactic Annotation-b Pos Tagging
Segmentation Level Word
Type Annotation Info
Annotation Type Alignment
Segmentation Level Sentence
Type Annotation Info
Annotation Type Lemmatization
Segmentation Level Word
Type Annotation Info
Character Encoding Info
Character Encoding UTF-8
Type Character Encoding Info
Language Info
Language pl
Language Polish
Language Name Polish
Size Per Language
Size 197,762,449
Size Unit Tokens
Type Size Info Type
Type Language Info
Language pt
Language Portuguese
Language Name Portuguese
Size Per Language
Size 211,824,204
Size Unit Tokens
Type Size Info Type
Type Language Info
Language sk
Language Slovak
Language Name Slovak
Size Per Language
Size 189,752,630
Size Unit Tokens
Type Size Info Type
Type Language Info
Language sq
Language Name Albanian
Size Per Language
Size 9,781,443
Size Unit Tokens
Type Size Info Type
Type Language Info
Language kk
Language Name Kazakh
Size Per Language
Size 486,766
Size Unit Tokens
Type Size Info Type
Type Language Info
Language he
Language Name Hebrew
Size Per Language
Size 2,872,765
Size Unit Tokens
Type Size Info Type
Type Language Info
Language tg
Language Name Tajik
Size Per Language
Size 160,123
Size Unit Tokens
Type Size Info Type
Type Language Info
Language hu
Language Name Hungarian
Size Per Language
Size 183,530,929
Size Unit Tokens
Type Size Info Type
Type Language Info
Language bs
Language Name Bosnian
Size Per Language
Size 6,195,646
Size Unit Tokens
Type Size Info Type
Type Language Info
Language lt
Language Name Lithuanian
Size Per Language
Size 170,381,570
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ga
Language Name Galician
Size Per Language
Size 629,272
Size Unit Tokens
Type Size Info Type
Type Language Info
Language es
Language Spanish
Language Name Spanish
Size Per Language
Size 191,092,782
Size Unit Tokens
Type Size Info Type
Type Language Info
Language mk
Language Name Macedonian
Size Per Language
Size 9,542,940
Size Unit Tokens
Type Size Info Type
Type Language Info
Language da
Language Danish
Language Name Danish
Size Per Language
Size 190,843,358
Size Unit Tokens
Type Size Info Type
Type Language Info
Language mn
Language Name Mongolian
Size Per Language
Size 135,076
Size Unit Tokens
Type Size Info Type
Type Language Info
Language is
Language Icelandic
Language Name Icelandic
Size Per Language
Size 762,894
Size Unit Tokens
Type Size Info Type
Type Language Info
Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#language Info
Language ar
Language Name Arabic
Size Per Language
Size 2,446,857
Size Unit Tokens
Type Size Info Type
Type Language Info
Language et
Language Name Estonian
Size Per Language
Size 160,175,247
Size Unit Tokens
Type Size Info Type
Type Language Info
Language sv
Language Swedish
Language Name Swedish
Size Per Language
Size 180,752,058
Size Unit Tokens
Type Size Info Type
Type Language Info
Language az
Language Name Azerbaijani
Size Per Language
Size 137,238
Size Unit Tokens
Type Size Info Type
Type Language Info
Language eu
Language Basque
Language Name Basque
Size Per Language
Size 461,080
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ro
Language Romanian
Language Name Romanian
Size Per Language
Size 235,859,637
Size Unit Tokens
Type Size Info Type
Type Language Info
Language tr
Language Name Turkish
Size Per Language
Size 13,297,328
Size Unit Tokens
Type Size Info Type
Type Language Info
Language hy
Language Name Armenian
Size Per Language
Size 139,802
Size Unit Tokens
Type Size Info Type
Type Language Info
Language tk
Language Name Turkmen
Size Per Language
Size 127,430
Size Unit Tokens
Type Size Info Type
Type Language Info
Language fr
Language French
Language Name French
Size Per Language
Size 231,486,663
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ru
Language Russian
Language Name Russian
Size Per Language
Size 3,293,243
Size Unit Tokens
Type Size Info Type
Type Language Info
Language uk
Language Name Ukrainian
Size Per Language
Size 744,815
Size Unit Tokens
Type Size Info Type
Type Language Info
Language hr
Language Name Croatian
Size Per Language
Size 11,950,183
Size Unit Tokens
Type Size Info Type
Type Language Info
Language mt
Language Name Maltese
Size Per Language
Size 163,515,445
Size Unit Tokens
Type Size Info Type
Type Language Info
Language fi
Language Finnish
Language Name Finnish
Size Per Language
Size 156,288,741
Size Unit Tokens
Type Size Info Type
Type Language Info
Language el
Language Greek (modern)
Language Name Greek
Size Per Language
Size 229,749,068
Size Unit Tokens
Type Size Info Type
Type Language Info
Language lv
Language Name Latvian
Size Per Language
Size 167,600,804
Size Unit Tokens
Type Size Info Type
Type Language Info
Language no
Language Norwegian
Language Name Norwegian
Size Per Language
Size 1,588,561
Size Unit Tokens
Type Size Info Type
Type Language Info
Language cs
Language Czech
Language Name Czech
Size Per Language
Size 196,769,297
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ja
Language Japanese
Language Name Japanese
Size Per Language
Size 50,194
Size Unit Tokens
Type Size Info Type
Type Language Info
Language zh
Language Chinese
Language Name Chinese
Size Per Language
Size 229,293
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ky
Language Name Kirghiz; Kyrgyz
Size Per Language
Size 135,031
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ka
Language Name Georgian
Size Per Language
Size 128,502
Size Unit Tokens
Type Size Info Type
Type Language Info
Language sl
Language Name Slovene
Size Per Language
Size 188,776,967
Size Unit Tokens
Type Size Info Type
Type Language Info
Language it
Language Italian
Language Name Italian
Size Per Language
Size 209,083,677
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ca
Language Catalan
Language Name Catalan; Valencian
Size Per Language
Size 640,522
Size Unit Tokens
Type Size Info Type
Type Language Info
Language de
Language German
Language Name German
Size Per Language
Size 194,497,872
Size Unit Tokens
Type Size Info Type
Type Language Info
Language ga
Language Name Irish
Size Per Language
Size 13,287,693
Size Unit Tokens
Type Size Info Type
Type Language Info
Language sr
Language Serbian
Language Name Serbian
Size Per Language
Size 1,832,323
Size Unit Tokens
Type Size Info Type
Type Language Info
Language nl
Language Dutch
Language Name Dutch
Size Per Language
Size 204,309,755
Size Unit Tokens
Type Size Info Type
Type Language Info
Linguality Info Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#linguality Info
Media Type Text
Size Info
Size 4,195,791,994
Size Unit Tokens
Type Size Info Type
Type Corpus Text Info
Resource Type Corpus
Type Corpus Info

Distribution Info

Availability Available-restricted Use
Availability Start Date 2011-09-01 Date
Ipr Holder
Organization Info
Communication Info
Address 52 Shipchenski prohod Blvd., Bl. 17
City Sofia
Country Bulgaria
Distribution Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#Dist URL4
Email [email protected]
Telephone Number +35 92 97 92 969
Type Communication Info
Zip Code 1113
Department Name Department of Computational Linguistics
Organization Name Institute for Bulgarian Language
Organization Short Name IBL
Type Organization Info Type
Type Actor
License
Delivery Channel Web Executable
Permission
Action http://creativecommons.org/ns/Distribution
http://creativecommons.org/ns/CommercialUse
Constraint Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#permission
Operator Eq
Purpose Academic Use
Type Prohibition
Constraint
Permission
Restrictions Of Use
Prohibition Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#permission
Same As Other
Type Licence Info
Delivery Channel Accessible Through Interface
Permission Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#permission
Prohibition Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#permission
Same As Other
Type Licence Info
Type Distribution
Distribution Info

Identification Info

Description The Bulgarian-X language Parallel Corpus (Bul-X-Cor) is a part of the Bulgarian National Corpus (BulNC). The Bulgarian National Corpus is designed as a uniform framework for texts of different modality (written - spoken), period (synchronic - diachronic), and number of languages (monolingual - parallel where one of the counterparts is Bulgarian). Any X-language in the corpus is equally treated with respect to the text type diversity and balance, metadata description scheme, preprocessing and annotation, search engine queries and data storage format. Bulgarian-X Language Parallel Corpus includes parallel corpora of 48 languages – English, German, French, Slavic and Balkan languages, as well as other European and non-European languages. The parallel corpora represent only texts which have a Bulgarian correspondence – either the original is in Bulgarian, there is a Bulgarian translation, or both texts are translations from a third language. As of January 2013, the Bulgarian-X Language Parallel Corpus contains 4.2 billion tokens, comprising the biggest parallel corpus of Bulgarian. Languages are not equally represented: the largest parallel corpus is the Bulgarian-English parallel corpus (280.8 and 283.1 million words for Bulgarian and English respectively); there are 18 other corpora of over 200 million tokens per language, 2 parallel corpora between 100 and 200 million tokens per language, 11 parallel corpora of size in the range 5-15 million tokens per language, and the rest 15 are below 1 million, with the smallest corpus being Japanese with 50,000 tokens. Each parallel subcorpus within Bul-X-Cor mirrors the structure of BulNC. The structure, data formatting and text description follow the model of BulNC. All Bulgarian texts in BulNC and English texts in Bul-X-Cor are supplied with extensive metadata description compliant with the well established standards. The Bulgarian-English parallel corpus is supplied as well with annotation on various levels while the annotation of other languages has just started. Main applications of parallel corpora are in the field of computational linguistics: machine translation, developing bilingual lexical resources (dictionaries), etc. The benefits of the parallel corpora increase if they are annotated. The Bulgarian-X Language Parallel Corpus Collocations service is a web service for collocations search and different types of statistics over the Bulgarian-X Language Parallel Corpus. The service employs the free of charge NoSketchEngine, a system for corpora processing that combines Manatee and Bonito. The Collocation service is a RESTful webservice, supporting complicated queries through http. Example: http://dcl.bas.bg/collocations/?cmd=collocations&word=нет user: bulnc pass: bulnc The query returns the collocations of a given word in the NoSketchEngine format. The system also supports additional arguments, namely all that are accepted by NoSketchEngine, provided with default values and an optional language identificator. The following example restricts the statistics to Bulgarian: http://dcl.bas.bg/collocations/?cmd=collocations&word=нет&lang=bg
Distribution
Access URL http://ibl.bas.bg/
http://search.dcl.bas.bg/en/
Type Distribution
URL
Access URL http://ibl.bas.bg/en/BGNC_en.htm
http://www.ibl.bas.bg/en/BGNC_parallel_en.htm
Type Distribution
URL
Identifier 805
Meta Share Id NOT_DEFINED_FOR_V2
Resource Short Name Bul-X-Cor
Title Bulgarian-X language Parallel Corpus
Type Identification Info

Resource Creation Info

Creator
Organization Info
Communication Info
Address 52 Shipchenski prohod Blvd., Bl. 17
City Sofia
Country Bulgaria
Distribution Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#Dist URL2
Email [email protected]
[email protected]
Fax Number +35 92 87 22 302
Telephone Number +35 92 97 92 969
+35 92 97 92 939
Type Communication Info
Zip Code 1113
Department Name Department of Computational Linguistics
Organization Name Institute for Bulgarian Language
Organization Short Name IBL
Type Organization Info Type
Type Actor
Funding Project
Distribution
Access URL http://cesar.nytud.hu/
Type Distribution
URL
Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#Dist URL
Funding Type Eu Funds
National Funds
Project End Date 2013-01-30 Date
2013-06-17 Date
Project Name Bulgarian National Corpus project
Central and South-East European Resources
Project Short Name CESAR
BulNC
Project Start Date 2009-12-17 Date
2011-02-01 Date
Type Project Info Type
Type Resource Creation Info

Resource Documentation Info

Documentation
Document Unstructured Koeva, Svetla, Ivelina Stoyanova, Svetlozara Leseva, Tsvetana Dimitrova, Rositsa Dekova, Ekaterina Tarpomanova. The Bulgarian National Corpus: Theory and Practice in Corpus Design. – Journal of Language Modelling, 2012, 1 (1), pp. 65-110. ISSN: 2299-8470 http://nlp.ipipan.waw.pl/ojs/index.php/JLM/issue/current
Type Documentation Info Type
Document Unstructured Koeva, Svetla, Ivelina Stoyanova, Rositsa Dekova, Borislav Rizov, Angel Genov. Bulgarian X-language Parallel Corpus. – In: Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12), Istanbul: European Language Resources Association (ELRA), 2012, pp. 51-62. ISBN: 978-2-9517408-7-7.
Type Documentation Info Type
Tool Documentation Type None
Manual
Help Functions
Type Resource Documentation Info

Usage Info

Actual Use Info
Actual Use Human Use
Type Actual Use Info
Foreseen Use Info
Foreseen Use Nlp Applications
Human Use
Type Foreseen Use Info
Type Usage Info

Validation Info

Type Validation Info
Validated true Boolean

Version Info

Has Version 2.0
Modified 2012-07-20 Date
Type Version Info

Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#Header

Instance of: Catalog Record
Issued 2014-09-23T00:41:58Z Date
Primary Topic Bulgarian-X language Parallel Corpus
Set Spec corpus:text
corpus

Metashare/b8ecf7fe66cd11e281b65cf3fcb88b70394683c3b32549349cf039716e61a92b#metadata Info

Instance of: Catalog Record
Created 2011-11-20 Date
Modified 2013-02-01 Date
Primary Topic Bulgarian-X language Parallel Corpus
Type Metadata Info