Index

Contact Point Metashare/8c13600ccd0711e1a404080027e73ea2f9cfd28f51d5437b8f5827c516c348fe#contact Person
Contributor Dan Tufis
Creator Amália Mendes
Description This lexicon includes multiword expressions (MWE) of European Portuguese extracted from a balanced 50,8M word written corpus – a subcorpus of the Reference Corpus of Contemporary Portuguese (CRPC). This corpus covers different genres, being mainly constituted by journalistic texts (59%), but it also includes texts from literature (21%), magazines (15%), miscellaneous, supreme court verdicts, parliament sessions and leaflets (5%). The MWE lexicon covers 1.198 lemmas (composed of single words from different POS categories: nouns, adjectives, verbs and adverbs) and a total of 12.753 MWE lemmas (which include inflectional variants of the MWE lemmas) and 242.233 concordances of those MWE expressions manually verified.
Rights underNegotiation
Source META-SHARE
Title LEX-MWE-PT: Word Combination in Portuguese Language
Type Lexical Conceptual Resource
Contact Point Metashare/12fdc090a35e11e1a404080027e73ea2c200dc17aff642fe980ba7a2da7f5ca1#contact Person
Contributor Dan Tufis
Creator Maria Fernanda Bacelar do Nascimento
Description This resource includes a spoken Portuguese corpus exemplifying the Portuguese spoken in Portugal, Brazil, Angola, Cape Verde, Guinea-Bissau, Mozambique, Sao Tome and Principe, Macao, Goa and East-Timor - with aligned sound and orthographic transcription - collected among sociolinguistically diverse speakers. It consists of recordings from informal conversations, conferences and media.
Rights underNegotiation
Source META-SHARE
Title Spoken Portuguese - Geographical and Social Varieties
Type Corpus
Contact Point Metashare/f30f4d04486111e2a2aa782bcb07413522a84532f1d443ffb62c5fcf3c59545a#contact Person
Contributor Dan Tufis
Creator Amália Mendes
Description The EUROPARL Corpus (subpart Portuguese-English of the parallel corpora), available at http://www.statmt.org/europarl/, was extracted from the proceedings of the European Parliament (Koehn, 2005). It contains transcriptions of sessions dating back from 1996 to 2011, in a total of approximately 58,324,562 tokens words of European Portuguese (L1) and 49,216,896 tokens of English (translation).
Language English
Portuguese
Rights CC-BY-SA
Source META-SHARE
Title EUROPARL Corpus Parallel Corpora: Portuguese-English
Type Corpus
Contact Point Metashare/2d875be6a35a11e1a404080027e73ea2fd1adfb79b3648ae9fff2932e82cf95a#contact Person
Contributor N/A
Creator N/A
Description This is the Maltese version of the Acquis Communautaire (AC), which is the total body of European Union (EU) law applicable in the EU Member States. It consists of selected texts between the 1950s and today, translated to Maltese.
Rights other
Source META-SHARE
Title Maltese Acquis Communautaire
Type Corpus
Contact Point Metashare/fe32ebf2485511e2a2aa782bcb074135aa0fdcd287ac45e7b67de9c36d8d2890#contact Person
Contributor Dan Tufis
Creator António Branco
Amália Mendes
Description CINTIL-Corpus Internacional do Português is a linguistically interpreted corpus of Portuguese. At present it is composed of 1 Million annotated tokens, verified by human expert annotators. The annotation comprises information on part-of-speech, open classes lemma and inflection, multi-word expressions pertaining to the class of adverbs and to the closed POS classes, and multi-word proper names (for named entity recognition). The corpus has been developed at the University of Lisbon by the NLX group at the Faculty of Sciences and the Anagrama group at the Cenro de Linguística da Universidade de Lisboa.
Language Portuguese
Rights ELRA_END_USER
Source META-SHARE
Title CINTIL-Corpus Internacional do Português
Type Corpus
Contact Point Metashare/2f2a00e4b92f11e1a404080027e73ea2eccd095ad8b0407989b2adb143ab6095#contact Person
Contributor Dan Tufis
Creator Rosa Del Gaudio
Description The corpus presented here is a collection of several tutorials and scientific papers in the field of Information Technology with 603 annotated definitions from Portuguese. The texts were collected from the Web at the beginning of the 2006 and they are organised in 32 files of three different sub-domains with 268,064 tokens: Information Society (91,825 tokens), Information Technology (80,483 tokens), and e-Learning (94,756 tokens).
Rights underNegotiation
Source META-SHARE
Title CINTIL-Definitions
Type Corpus
Contact Point Metashare/72cc03d88be311e294080015171445924c5e7298251340b291958854f48d783e#contact Person
Metashare/72cc03d88be311e294080015171445924c5e7298251340b291958854f48d783e#contact Person2
Contributor Anđelka Zečević
Creator Anđelka Zečević
Krstev Cvetana
Description NERosetta is a multiuser web application that aims to facilitate retrieval and comparison of named entities in a single or parallel texts. The main named entity categorization is realized according to the Quaero annotation recommendation and provides a user with approximately 50 different search options. Registered users have an extra possibility to share annotated resources (in XML format) and annotation schemas for the set of languages as well as to manage their own resources and schemas. The initial version supports four annotation schemas (Stanford NER 3 and Stanford NER 7 for English, Krstev&Vitas for Serbian and Maurel for French) and three annotated parallel versions of Jules Verne's Around The World in Eighty Days (English-Serbian, French-Serbian and French-English).
Rights GPL
Source META-SHARE
Title NERosetta
Type Tool Service
Contact Point Metashare/362a2020cf5711e1a404080027e73ea28eaaf998e9aa47739841451ea4e16f51#contact Person
Contributor Dan Tufis
Creator Maria Fernanda Bacelar do Nascimento
Description This resource includes a spoken corpus with approximately 300.000 words, covering both formal (152.755 words) and informal (165.838 words) speech, with aligned sound and orthographic transcription and POS-tag information.
Rights ELRA_END_USER
Source META-SHARE
Title C-ORAL-ROM_EXM
Type Corpus
Contact Point Metashare/27607ab28b2c11e2975a00151714459237c30a10120c409ab292a1ed8f3ec9fc#contact Person2
Metashare/27607ab28b2c11e2975a00151714459237c30a10120c409ab292a1ed8f3ec9fc#contact Person
Contributor Mirko Spasić
Creator Mirko Spasić
Duško Vitas
Description Serbian NGrams (SrpNGrams) represent set of N-grams extracted from Serbian Lemmatized and PoS Annotated Corpus (SrpLemKor) for N from 1 to 5. Each unigram is maximum continuous chunk of non-whitespace lower-case characters. The resource contains all unique N-grams preceded by number of occurrencies. It also contains n-gram language models (1-5) in the standard ARPA text and binary format, created by IRST Language Modeling Toolkit. SrpKor texts consist of: fiction written by Serbian authors in 20th and 21th century, various scientific texts from various domains (both humanities and sciences), legislative texts and general texts. General texts represent daily news published in newspaper \"Politika\" 2000-2002 and 2005-2010, texts in journals and magazines 1991-2002 (\"Danica\", \"Ebit\", \"Ekonomist\", \"Glasnik\", \"NIN\", \"Ilustrovana politika\", \"Kalibar\", \"Moje srce\", \"Mostovi\", \"Pravoslavlje\", \"Svet\", \"Teološki pogledi\", \"Trn\", \"Viva\", \"Republika\"), internet portal texts 2011-2012 (Peščanik), TANJUG agency news 1995-96, newspaper feuilletons published in newspapers \"Politika\" (2001-2003), \"Večernje novosti\" (2008-2011) and \"Danas\" (2002-2006).
Rights MS-NC-NoReD-ND
Source META-SHARE
Title Serbian NGrams
Type Corpus
Contact Point Metashare/58b341b48be311e294150015171445929ea8576278db448b817323b587beaf95#contact Person
Metashare/58b341b48be311e294150015171445929ea8576278db448b817323b587beaf95#contact Person2
Contributor Miljana Mladenović
Creator Cvetana Krstev
Miljana Mladenović
Description This tool is a web application for ontological based emotions recognition and tagging of Serbian texts. The application uses RDFS which are created by using nine discrete emotion psychological theories. Also, it uses associative dictionary of Serbian with about 11 thousands words and Serbian morphological electronic dictionary which contains approximately 4.4 million different inflectional forms of simple words. The application offers a representation of summary results in a graphical form. Annotation of an uploaded text is possible for XML and textual documents as well as a text from Web.
Rights GPL
Source META-SHARE
Title Emotions Annotation Tool
Type Tool Service
Contact Point Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person2
Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person
Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person3
Contributor Rui Lageira
Gonçalo Simões
Helena Galhardas
Creator Gonçalo Simões
Description Etxt2DB is a framework for specifying and executing Entity Recognition (ER) programs. These programs accept as input a text containing potentially interesting entities to be extracted and produce the input text annotated with the recognized entities. The Etxt2DB functioning mode involves two distinct phases. First, the training phase consists in creating a model based on a given ER technique and one or more resources that guide the creation of the classification model. Examples of these resources are dictionaries for rule-based ER techniques or training data for statistical learning techniques (e.g., Conditional Random Fields). Second, in the execution phase, a classification model previously created receives as input plain text and produces annotations corresponding to the recognized entities. The Etxt2DB framework consists of a software layer, built on top of Minorthird and Lingpipe, offering a command-like specification language. Existing Machine Learning Java APIs (such as Minorthird and Lingpipe) provide implementations of Entity Recognition techniques. Some developers of ER applications do not want to get involved in the implementation details of the techniques used. Instead, they are willing to focus on: the choice of the technique to be used; the resources used in the process (e.g., dictionaries); a good set of features that help the ER program to take adequate decisions. The objective of the Etxt2DB specification language is to turn the development and tuning of ER programs easier for developers that are mainly concerned with these topics. In the context of the METANET project, the goal was to build a component-generator tool that encapsulates Etxt2DB. In the training phase, this tool accepts a training data set as input and produces a classification model and a U-Compare component that is able to interpret that model. In the execution phase, the component produced is loaded into the U-Compare platform and then is ready to be used for recognizing entities from text.
Rights GPL
Source META-SHARE
Title U-Compare E-txt2DB: Giving structure to unstructured data
Type Tool Service
Contact Point Metashare/0cb6205066e111e2bac9525400d761476ef3f57fa58942a4aac54d4206216320#contact Person
Contributor Łukasz Dróżdż
Piotr Pęzik
Creator Łukasz Dróżdż
Piotr Pęzik
Description A subset of the PELCRA PLEC corpus, containing 15 hours (131 000 transcribed words) of recordings of informal interviews with Polish learners of English, time-aligned on the utterance and annotated manually for mispronounciation errors, provided as TEI P5-conformant XML and EAF (ELAN) files.
Language English
Rights CC-BY-NC
Source META-SHARE
Title PELCRA Spoken Learner English Corpus
Type Corpus
Contact Point Metashare/b6646fb866e011e29895525400d761474918098ee988487eaf619d3da163a80b#contact Person2
Metashare/b6646fb866e011e29895525400d761474918098ee988487eaf619d3da163a80b#contact Person
Contributor Piotr Pęzik
Łukasz Dróżdż
Creator Łukasz Dróżdż
Piotr Pęzik
Description A subset of the PELCRA corpus of conversational Polish, time-aligned on the utterance level, licensed under the CC-BY-NC license. This resource contains 386 744 words in 73 transcriptions of over 43 hours of recordings made in the years 2008-2010. The texts are provided as TEI P5-compliant XML files with custom PELCRA extensions and in the XLIFF format.
Language Polish
Rights CC-BY-NC
Source META-SHARE
Title PELCRA time-aligned spoken corpus of Polish (CC-BY-NC)
Type Corpus
Contact Point Metashare/1f82f6866b0011e284b6000423bfd61c95584808dd944c28a354bba3bb390dbd#contact Person
Contributor Alina Wróblewska
Creator Alina Wróblewska
Description Statistical dependency parsing model is trained on the Polish Dependency Bank (PDB, Pol. Składnica zależnościowa) with the the publicly available parsing system -- MaltParser. MaltParser is a transition-based dependency parser that uses a deterministic parsing algorithm. The deterministic parsing algorithm builds a dependency structure of an input sentence based on transitions (shift-reduce actions) predicted by a classifier. The classifier learns to predict the next transition given training data and the parse history.
Rights GPL
Source META-SHARE
Title Dependency Parsing Model for Polish
Type Tool Service
Contact Point Metashare/4afd693e6ba711e2aa7c68b599c26a0651325e2993be4cc2a4950546705978a0#contact Person2
Metashare/4afd693e6ba711e2aa7c68b599c26a0651325e2993be4cc2a4950546705978a0#contact Person
Contributor Attila Mártonfi
Description Hungarian historical corpus (further as HHC) is a collection of texts written between 1772 and 1997 in different genres, containing ca. 27 million tokens. During the compilation of HHC, text samples were selected by professionals (literary historians, historians, mathematicians etc.) from printed works. A relative majority (40%) of the texts are dated from the second half of the 20th century. The corpus is the product of the Department of Lexicography and Lexicology at RIL HAS, made between 1986 and 1997, maintained continuously since then. As an innovation, genre labeling was unified. Thus, genres and text types in HHC and HNC are marked similarly, this makes possible to search data of these corpora by using the same query structure.
Rights MS-NC-NoReD
Source META-SHARE
Title HHC: Hungarian historical corpus
Type Corpus
Contact Point Metashare/dad2b9848be011e29ebd001517144592d5a00254a9fd45bb9383caa72801461a#contact Person
Contributor Cvetana Krstev
Creator Cvetana Krstev
Description Morphological electronic dictionary of Serbian (Ekavian pronunciation) (SrpMD) released in the scope of the EU-funded CESAR project is a version of morphological dictionary of Serbian used in the NooJ corpus processing system and consituting the part of the Serbian Nooj Module (see section 6.8). This version is compliant to MULTEXT-East morphosyntactic specification for Serbian (http://nl.ijs.si/ME/V4/msd/html/msd-sr.html) (with one small deviation form it – see section 6.10). It comprises of 3,630,613 entries for 85,721 lemmas covering 11 PoS: nouns (646,867/40,425), adjectives (2,315,640/25,826), verbs (654,159/15,359), adverbs (3233), numerals(4,794/175), conjunctions (83), interjections (218), prepositions (169), pronouns (5,321/104), particles (103), abbreviations (26).
Rights MS-NC-NoReD
Source META-SHARE
Title Serbian Morphological Dictionary (Multext-East)
Type Lexical Conceptual Resource
Contact Point Metashare/936f54fe8bdf11e2bffb0015171445921ba762cd361b47939ea5c40cb0b79fbe#contact Person
Contributor Miloš Utvić
Creator Miloš Utvić
Ivan Obradović
Duško Vitas
Description This corpus consists of English source texts translated into Serbian, and Serbian source texts translated into English, and several aligned English and Serbian translations of literary texts originally in French. The texts belong to various domains: fiction, general news, scientific journals, web journalism, health, law, education, movie sub-titles. The corpus also contains several Serbian translations of texts from the ‘Acquis communautaire’ corpus and from the ‘Intera’ corpus aligned with their originals. The alignment was performed on the subsentencial level. The texts were segmented and aligned automatically and then manually checked. In most cases the alignment is one-to-one. The size of the corpus is 5,078,280 words (2,672,911 in the English part, 2,405,369 in the Serbian part). More about the content of this corpus can be found at: http://www.korpus.matf.bg.ac.rs/SrpEngKor/SrpEngKor_2013_01.pdf
Language English
Rights CC-BY-NC
Source META-SHARE
Title English-Serbian Aligned Corpus
Type Corpus
Contact Point Metashare/0d68b2f28b3411e2ab9f001517144592e9978ff1de0d4abebd4d6c8935fcb9af#contact Person
Contributor Miloš Utvić
Ranka Stanović
Creator Miloš Utvić
Duško Vitas
Ranka Stanković
Description Serbian NooJ module (SrpNooJ) was produced in the scope of the EU-funded CESAR project. It consists of a set of resources in both alphabets that are in use for Serbian: Cyrillic and Latin. Each set consists of: the dictionary properties’ definition file (metadata), one text – a novel “Dva carstva” (Two empires) from a Serbian author Branimir Ćosić comprising of 106684 tokens, a sample dictionary in readable form with 35 lemma that belong to 9 grammatical classes, with examples of multiword units and derivational morphology, a sample of morphological grammars used for lemmas from a sample dictionary – three for simple nouns, two for adjectives, two for verbs, and one for a multiunit noun, a readable sample dictionary of inflected forms automatically produced from a sample dictionary of lemmas and a sample morphological grammars, a syntactic grammar for recognition of one class of named entities – full personal names with their roles or functions, a full compiled dictionary (divided in three files: nouns, verbs, and other). It comprises of 85868 entries: nouns (40886), adjectives (25558), verbs (15366), and other (4058).
Rights CC-BY
Source META-SHARE
Title Serbian NooJ module
Type Lexical Conceptual Resource
Contact Point Metashare/fc91787a6b7f11e29f6e000423bfd61cad17bb05bcbd470da8cec4ebdda3481e#contact Person
Contributor Max Silberztein
Creator Mladen Stanojević
Description NooJ is a linguistic development environment that allows linguists to formalize several levels of linguistic phenomena: typography and spelling; lexicons of simple words, multiword units and discontinuous expressions; inflectional, derivational and productive morphology; local and structural syntax, transformational and semantic analysis and generation. For each of these levels NooJ provides linguists with one formal framework specifically designed to facilitate the description of each phenomenon, as well as parsing/development/debugging tools designed to be as computationally efficient as possible, from Finite-State machines to Turing machines. This approach distinguishes NooJ from other computational linguistic frameworks which provide a unique formalism based on a compromise between power and efficiency. As a corpus processing tool, NooJ allows all researchers and professional to extract information from general or technical corpora by applying sophisticated queries based on concepts rather than word forms and build indices, add semantic annotations, perform statistical analyses, etc. MONO version of NooJ is operative on all platforms that support MONO.
Rights MS-NC-NoReD-ND
Source META-SHARE
Title MONO version of NooJ
Type Tool Service