Index

Contact Point Metashare/aee9fa36de6911e2b1e400259011f6ea9a9fa363e71e4076846ac55a434fff9e#contact Person
Description Cette base de données comprend les modèles HTS Festival bilingues (anglais et espagnol). Les modèles ont été entraînés à partir de 9 heures de parole réalisés par 2 locuteurs bilingues femmes et 2 locuteurs bilingues hommes. Chaque locuteur a enregistré 2h 15 min par langue. La base de données de parole peut être trouvée dans la base de données orale de conversion vocale bilingue TC-STAR pour l’espagnol (ELRA-S0311) et dans la base de données de parole expressive bilingue TC-STAR (ELRA-S0313).
This database contains Bilingual (English and Spanish) Festival HTS models. Models were trained with 9h of speech from 2 female bilingual speakers and 2 male bilingual speakers. Each speaker recorded 2h 15 min per language. The speech data can be found in the TC-STAR Bilingual Voice-Conversion Spanish Speech Database (ELRA-S0311) and in the TC-STAR Bilingual Expressive Spanish Speech Database (ELRA-S0313).
Language Spanish
English
Rights ELRA_END_USER
ELRA_VAR
Source META-SHARE
Title Bilingual (Spanish-English) Speech synthesis HTS models
Modèles HTS bilingues pour la synthèse vocale (espagnol-anglais)
Type Corpus
Contact Point Metashare/0292790ade6b11e2b1e400259011f6ea65e04fe27a1d42188fd828ea1257aede#contact Person
Description In 1996, some 75 Dutch people participated in recording a multi-purpose continuous speech database. Most of them were recruited from the TNO Human Factors Research Institute, where the recordings were made. The main part of the database consisted of Dutch sentences. However, most speakers participated in recording 10 sentences in English, French and German. This data was initially distributed as a common data set for research leading to presentations and discussions at the ESCA/NATO MIST workshop held in Leusen, The Netherlands, in 1999. The non-nativeness in any particular language, for instance English, is of course very biased towards Dutch, and therefore this database can be considered only as a start for studying non-native speech. However, with experiences with this database, researchers in other countries may record similar data, so that also other foreign accents can be studied, and compared to this database. Recording conditions: - Sennheiser HMD-414-6 close talking microphone - B&K MD-211-N far-field microphone - anechoic silent recording room - sentences read from computer screen - Ariel Pro-Port digital recording equipment - 16 kHz sampling rate, 16 bit resolution Speech material - 10 sentences in Dutch, English, French and German, including 5 sentences per language which are identical for all speakers and 5 sentences per language which are unique for each speaker - Sentence text from newspapers: Dutch: NRC/Handelsblad; English: Wall Street Journal; French: Le Monde; German: Frankfurter Rundschau The text of the English, French and German sentences were obtained from other databases recorded/used in the European project ‘SQALE’. Annotation: - Dutch sentences are orthographically annotated - For English, French and German sentences the prompt texts are available - Only the Dutch unique sentences have been listened to, and annotated accordingly. The English, French and German sentences have been generated from the prompt texts, i.e., only the punctuation characters have been removed. For French and English, the first word has been de-capitalized according to some simple algorithm. - The spoken text is annotated in a format of one line per speech utterance, with the utterance identification in parenthesis at the end. Speakers: - 74 speakers, including 52 males and 22 females - All speakers are native Dutch. Not all of them were able to produce speech in German, English and French.
En 1996, 75 locuteurs hollandais ont participé à l’enregistrement d’une base de données de parole continue multi-objectifs. La plupart d’entre eux ont été recrutés par L’institut de recherche sur les facteurs humains de TNO, où les enregistrements ont été réalisés. La plus grande partie de la base de données consistait en des phrases en hollandais. Cependant, la plupart des locuteurs ont également participé à l’enregistrement de 10 phrases en anglais, en français et en allemand. Ces données ont d’abord été distribuées sous la forme d’un ensemble de données communes pour la recherche qui a conduit à des présentations et des discussions lors de l’atelier ESCA/NATO MIST, de Leusen, aux Pays-Bas, en 1999. Le fait d’être locuteur non natif d’une langue, par exemple l’anglais, est bien sûr très biaisé vis-à-vis du hollandais, et cette base de données peut donc ainsi être considérée uniquement comme une base initiale pour l’étude de la parole non native. Cependant, grâce aux expériences réalisées avec cette base, les chercheurs d’autres pays peuvent enregistrer des données similaires, afin que d’autres accents étrangers puissent être étudiés et être comparés à cette base. Conditions d’enregistrements : - Micro-casque Sennheiser HMD-414-6 - Microphone placé à distance (“far-field”) B&K MD-211-N - enregistrement en chambre sourde - phrases lues sur écran d’ordinateur - équipement d’enregistrement numérique Ariel Pro-Port - taux d’échantillonnage de 16 kHz, résolution de 16 bit Matériel de parole : - 10 phrases en hollandais, anglais, français et allemand, dont 5 phrases identiques par langue pour tous les locuteurs et 5 phrases distinctes par langue et par locuteur - Phrases extraites de journaux: NRC/Handelsblad pour le hollandais, Wall Street Journal pour l’anglais, Le Monde pour le français, Frankfurter Rundschau pour l’allemand Le texte des phrases en anglais, français et allemand a été obtenu à partir d’autres bases de données enregistrées/utilisées dans le projet européen « SQALE ». Annotation : - Les phrases en hollandais sont annotées au niveau orthographique - Pour les phrases en anglais, français et allemand, les textes énoncés sont disponibles - Seules les phrases distinctes en hollandais ont été écoutées et annotées. Les phrases en anglais, français et allemand ont été générées à partir des textes énoncés, c’est-à-dire que seuls les caractères de ponctuation ont été supprimés. Pour le français et l’anglais, la majuscule du premier mot a été supprimée grâce à un algorithme simple. - Le texte parlé est annoté au format suivant : une ligne par occurrence de parole, avec l’identification de l’occurrence entre parenthèses à la fin. Locuteurs : - 74 locuteurs, dont 52 hommes et 22 femmes - Tous les locuteurs sont natifs du hollandais. Tous n’étaient pas capables de produire de la parole en allemand, anglais et français.
Language English
Rights ELRA_END_USER
Source META-SHARE
Title MIST Multi-lingual Interoperability in Speech Technology database
Base de données MIST (Multi-lingual Interoperability in Speech Technology)
Type Corpus
Contact Point Metashare/fd11707cde7311e2b1e400259011f6ea85d38ca6430c4730909ac9faf64ee2a6#contact Person
Description * Entrées anglais-espagnol : Recherche scientifique & sciences mathématiques (906 entrées), géosciences (10 215), informatique, électronique & télécommunications (70 580), industrie (47 578), transports & maintenance (12 291), économie (145 572), sciences biologiques (38 989), communication & média (8 143), sciences chimiques & physiques (27 467). * Entrées allemand-anglais-espagnol-français : Environnement (36 658), santé (66 727), agriculture & alimentation (25 975), construction & travaux publics (8 429), droit & politique (56 578), sports & loisirs (17 312). * Deux lexiques spécialisés: Espagnol-anglais et allemand-anglais-français sans codes de domaine : électronique, télématique, droit, taxes, douanes, etc. (550 000 entrées). * Deux lexiques généraux: Allemand-anglais-espagnol-français et allemand-anglais-espagnol-français-italien-portugais. (83 000 entrées). Cette base de données terminologique contient, pour chaque domaine, l'indication de sous-domaines (de 2 sous-domaines pour la recherche scientifique à 39 pour les sports et loisirs). Chaque entrée comporte une définition, une unité phraséologique, une abréviation, une information sur l'usage et des étiquettes grammaticales. Format: ASCII Support : disquette
* Entries for English-Spanish: Scientific research & mathematical sciences (906 entries), Geosciences (10,215), Computer science, electronics & telecommunications (70,580), Industry (47,578), Transport & Maintenance (12,291), Economy (145,572), Biological sciences (38,989), Communication & media (8,143), Chemical & physical sciences (27,467). * Entries for English-French-German-Spanish: Environment (36,658), Health (66,727), Agriculture & food (25,975), Construction & public works (8,429), Law & policy (56,578), Sports & Leisure (17,312) * Two specialized lexicons: Spanish-English and English-French-German without domain codes: electronics, telematics, law, taxes, customs, etc. (550,000 entries). * Two general lexicons: Spanish-English-French-German and Spanish-English-French-German-Portuguese-Italian (83,000 entries). This terminological database contains, for each domain, a sub-domain indication is given (from 2 sub-domains for Scientific research to 39 for Sports & leisure). Each entry consists of a definition, phraseological unit, abbreviation, usage information, grammatical labels. Format: ASCII Medium: floppy disk
Language English
Spanish
Rights ELRA_VAR
Source META-SHARE
Title Base de données terminologique polytechnique et plurilingue VERBA - D-AE Contrôle climatique
VERBA Polytechnic and Plurilingual Terminological Database - D-AE Climate Control
Type Lexical Conceptual Resource
Contact Point Metashare/f9be99cabbb611e28763000c291ecfc8c5698cb825a64a5f91cc4a4866705914#contact Person
Description This data set contains Spanish word n-grams and Spanish word/tag/lemma n-grams in the \"Environment\" (ENV) domain. N-grams are accompanied by their observed frequency counts. The length of the n-grams ranges from unigrams (single words) to five-grams. The data were collected in the context of PANACEA (http://www.panacea-lr.eu), an EU-FP7 Funded Project under Grant Agreement 248064. The n-gram counts were generated from crawled Web pages that were automatically detected to be in the Spanish language and were automatically classified as relevant to the ENV domain. The ENV domain collection used consisted of approximately 49.86 million tokens. Data collection took place in the summer of 2011.
Language Spanish
Rights CC-BY-SA
Source META-SHARE
Title PANACEA Environment Corpus n-grams ES (Spanish)
Type Corpus
Contact Point Metashare/99a27130de7311e2b1e400259011f6ea02c79f0d57784efeb067a83f385bc98c#contact Person
Description The Aurora project was originally set up to establish a world wide standard for the feature extraction software which forms the core of the front-end of a DSR (Distributed Speech Recognition) system. ETSI formally adopted this activity as work items 007 and 008.The two work items within ETSI are : - ETSI DES/STQ WI007 : Distributed Speech Recognition - Front-End Feature Extraction Algorithm & Compression Algorithm - ETSI DES/STQ WI008 : Distributed Speech Recognition - Advanced Feature Extraction Algorithm. This database is a subset of the SpeechDat-Car database in Danish language which has been collected as part of the European Union funded SpeechDat-Car project. It contains isolated and connected Danish digits spoken in the following noise and driving conditions inside a car : 1. High speed good road 2. Low speed rough road 3. Stopped with motor running 4. Town traffic
DESCRIPTION DISPONIBLE EN FRANCAIS PROCHAINEMENT. The Aurora project was originally set up to establish a world wide standard for the feature extraction software which forms the core of the front-end of a DSR (Distributed Speech Recognition) system. ETSI formally adopted this activity as work items 007 and 008.The two work items within ETSI are : - ETSI DES/STQ WI007 : Distributed Speech Recognition - Front-End Feature Extraction Algorithm & Compression Algorithm - ETSI DES/STQ WI008 : Distributed Speech Recognition - Advanced Feature Extraction Algorithm. This database is a subset of the SpeechDat-Car database in Danish language which has been collected as part of the European Union funded SpeechDat-Car project. It contains isolated and connected Danish digits spoken in the following noise and driving conditions inside a car : 1. High speed good road 2. Low speed rough road 3. Stopped with motor running 4. Town traffic
Language Dnj
Rights ELRA_END_USER
Source META-SHARE
Title AURORA Project database - Subset of SpeechDat-Car - Danish database - Evaluation Package
Base de données du projet AURORA - sous-ensemble de la base de données SpeechDat-Car du danois - Package d'évaluation
Type Corpus
Contact Point Metashare/c366848692c211e28763000c291ecfc8720a7e22a70f48ec960d5887b7e4a007#contact Person2
Metashare/c366848692c211e28763000c291ecfc8720a7e22a70f48ec960d5887b7e4a007#contact Person
Creator Jimmy O'Reagan
Description This is the LMF version of the Apertium bilingual dictionary for French and Catalan languags. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as English-Catalan). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Language Catalan
French
Rights GPL
Source META-SHARE
Title French-Catalan LMF Apertium Bilingual dictionary
Type Lexical Conceptual Resource
Contact Point Metashare/ef504ebede7211e2b1e400259011f6ea00426fc7d67744068f128b774af09a9b#contact Person
Description * Entrées anglais-espagnol : Recherche scientifique & sciences mathématiques (906 entrées), géosciences (10 215), informatique, électronique & télécommunications (70 580), industrie (47 578), transports & maintenance (12 291), économie (145 572), sciences biologiques (38 989), communication & média (8 143), sciences chimiques & physiques (27 467). * Entrées allemand-anglais-espagnol-français : Environnement (36 658), santé (66 727), agriculture & alimentation (25 975), construction & travaux publics (8 429), droit & politique (56 578), sports & loisirs (17 312). * Deux lexiques spécialisés: Espagnol-anglais et allemand-anglais-français sans codes de domaine : électronique, télématique, droit, taxes, douanes, etc. (550 000 entrées). * Deux lexiques généraux: Allemand-anglais-espagnol-français et allemand-anglais-espagnol-français-italien-portugais. (83 000 entrées). Cette base de données terminologique contient, pour chaque domaine, l'indication de sous-domaines (de 2 sous-domaines pour la recherche scientifique à 39 pour les sports et loisirs). Chaque entrée comporte une définition, une unité phraséologique, une abréviation, une information sur l'usage et des étiquettes grammaticales. Format: ASCII Support : disquette
* Entries for English-Spanish: Scientific research & mathematical sciences (906 entries), Geosciences (10,215), Computer science, electronics & telecommunications (70,580), Industry (47,578), Transport & Maintenance (12,291), Economy (145,572), Biological sciences (38,989), Communication & media (8,143), Chemical & physical sciences (27,467). * Entries for English-French-German-Spanish: Environment (36,658), Health (66,727), Agriculture & food (25,975), Construction & public works (8,429), Law & policy (56,578), Sports & Leisure (17,312) * Two specialized lexicons: Spanish-English and English-French-German without domain codes: electronics, telematics, law, taxes, customs, etc. (550,000 entries). * Two general lexicons: Spanish-English-French-German and Spanish-English-French-German-Portuguese-Italian (83,000 entries). This terminological database contains, for each domain, a sub-domain indication is given (from 2 sub-domains for Scientific research to 39 for Sports & leisure). Each entry consists of a definition, phraseological unit, abbreviation, usage information, grammatical labels. Format: ASCII Medium: floppy disk
Language English
Spanish
Rights ELRA_VAR
Source META-SHARE
Title VERBA Polytechnic and Plurilingual Terminological Database - G-GR Ionics
Base de données terminologique polytechnique et plurilingue VERBA - G-GR Physique ionique
Type Lexical Conceptual Resource
Contact Point Metashare/c855da065a6811e29a5400504503039c72725b558f9b4f1bae4c5a3b5ac39aa8#contact Person
Description A corpus with texts from Göteborgsposten
En korpus med texter från Göteborgsposten
Language Swedish
Rights CC-BY
other
Source META-SHARE
Title GP 2010
GP 2010
Type Corpus
Contact Point Metashare/c97bdcfe92c211e28763000c291ecfc80a826828b499489f9b17b0aff28ee998#contact Person2
Metashare/c97bdcfe92c211e28763000c291ecfc80a826828b499489f9b17b0aff28ee998#contact Person
Description This is the LMF version of the Spanish Freeling Sense. FreeLing is a developer-oriented library providing language analysis services. FreeLing is designed to be used as an external library from any application requiring this kind of services. Nevertheless, a simple main program is also provided as a basic interface to the library, which enables the user to analyze text files from the command line. The original Catalan and Spanish sense dictionaries are extracted from EuroWordNet, and the reduced subsets included in this FreeLing package are distibuted under GNU GPL license.
Language Spanish
Rights GPL
Source META-SHARE
Title Spanish LMF Freeling Sense
Type Lexical Conceptual Resource
Contact Point Metashare/f73b18cc5a6811e29a5400504503039ca4aa15146248480b8e13540df445473b#contact Person
Description Texter från Dramawebben, ett digitalt arkiv över fri svensk dramatik.
Texts from Dramawebben, a digital archive of free Swedish drama.
Language Swedish
Rights CC-BY
other
Source META-SHARE
Title Dramawebben (demo)
Dramawebben (demo)
Type Corpus
Contact Point Metashare/93c9c4f4de6c11e2b1e400259011f6ead57b054cfad346fe995d9c5728f01c46#contact Person
Description Contrats d'assurance, assurance publique et privée, ressources terminologiques utilisées dans les institutions de l'Union Européenne. Fiches disponibles : 1000 Langues : Catalan, Espagnol, Anglais Format: ASCII Support : disquette Description de la fiche : Chaque fiche de cette base terminologique contient une définition, des abréviations, des notes, des étiquettes grammaticales (catégorie, genre et nombre), synonymes.
Insurance contracts, private and public insurance, resource terminology used within European Union institutions. Cards available: 1000 Languages: Catalan, Spanish, English Format: ASCII Medium: floppy disk Card Description: Each card in this terminological database contains a definition, abbreviations, notes, grammatical labels (category, gender and number), synonyms.
Language Catalan
Spanish
English
Rights ELRA_VAR
ELRA_END_USER
Source META-SHARE
Title Assurance (Termcat)
Insurance (Termcat)
Type Lexical Conceptual Resource
Contact Point Metashare/bcd5238492c211e28763000c291ecfc8ad3bef674daf4e53abfe5f0ab2061764#contact Person
Metashare/bcd5238492c211e28763000c291ecfc8ad3bef674daf4e53abfe5f0ab2061764#contact Person2
Creator Paul Breen
Jimmy O'Reagan
Description This is the LMF version of the Apertium bilingual dictionary for English and Catalan languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as English-Catalan). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Language English
Catalan
Rights GPL
Source META-SHARE
Title English-Catalan LMF Apertium Bilingual dictionary
Type Lexical Conceptual Resource
Contact Point Metashare/7788c0c2de6e11e2b1e400259011f6ea9ebbf368dfbf4c7ab069e5edbe1bd010#contact Person
Description Le lexique phonétique LC-STAR espagnol a été créé dans le cadre du projet LC-STAR (IST 2001-32216), financé par la Commission européenne et le gouvernement espagnol. Le lexique a été produit au Centre de technologies et d’applications de la langue et de la parole (TALP) de l’Universitat Politècnica de Catalunya (UPC) (Espagnol), qui en est également le propriétaire. Le lexique comprend plus de 100 000 mots répartis en trois catégories : - Un ensemble de 55 854 mots communs. Cet ensemble a été extrait d’un corpus de plus de 20 millions de mots répartis en 6 domaines différents (sports/jeux, actualités, finances, culture/amusement, information consommateur, communications personnelles), avec pour objectif d’atteindre pour chaque domaine au moins 95% de couverture. En plus des listes de mots extraites du corpus, une liste de classes de mots en ensemble fermé (fonctions) est incluse dans la liste de mots finale. - Un ensemble de 45 403 noms propres (noms de personnes, noms de familles, villes, rues, sociétés, noms de marque) divisée en 3 domaines. Les noms comportant des mots multiples, tels que New_York, ont été conservés dans chacun des 3 domaines et comptent ainsi pour une seule entrée. Les 3 domaines sont : prénoms et noms de familles (23 114 entrées différentes), noms de lieux (15 427 entrées différentes), et organisations (7 777 entrées différentes). - Une liste de 7 498 mots d’application traduits à partir de termes anglais tels que définis par le consortium LC-STAR. Cette liste comprend des nombres, des lettres, des abréviations et du vocabulaire spécifique aux applications contrôlées par la voix (recherche d’information, contrôle des appareils de consommation, etc.). Le lexique est fourni au format XML et inclut des transcriptions phonétiques en SAMPA. La base de données est stockée sur 1 CD.
The LC-STAR Spanish phonetic lexicon was created within the scope of the LC-STAR project (IST 2001-32216) which was sponsored by the European Commission and the Spanish Government. Production was performed at the Technologies and Applications of Language and Speech Center (TALP) of the Universitat Politècnica de Catalunya (UPC) (Spain). The owner of the database is UPC. The lexicon comprises more than 100,000 words, distributed over three categories: - a set of 55,854 common word entries. This set is extracted from a corpus of more than 37 million words distributed over 6 different domains (sports/games, news, finance, culture/entertainment, consumer information, personal communications). This was done with the aim of reaching a target for each domain of at least 95% self coverage. In addition to extracting word lists from the corpus, a list of closed set (function) word classes are included in the final word list. - a set of 45,403 proper names (including person names, family names, cities, streets, companies and brand names) divided into 3 domains. Multiple word names such as New_York are kept together in all three domains, and they count as one entry. The 3 domains consist of first and last names (23,114 different entries), place names (15,427 different entries), and organisations (7,777 different entries). - and a list of 7,498 special application words translated from English terms defined by the LC-STAR consortium. This list contains: numbers, letters, abbreviations and specific vocabulary for applications controlled by voice (information retrieval, controlling of consumer devices, etc.). The lexicon is provided in XML format and includes phonetic transcriptions in SAMPA. The database is stored on 1 CD.
Language Spanish
Rights ELRA_VAR
ELRA_END_USER
Source META-SHARE
Title LC-STAR Spanish phonetic lexicon
Lexique phonétique LC-STAR espagnol
Type Lexical Conceptual Resource
Contact Point Metashare/e2bc95b492c211e28763000c291ecfc8492939bf6d024f99932fa452e636dce4#contact Person
Metashare/e2bc95b492c211e28763000c291ecfc8492939bf6d024f99932fa452e636dce4#contact Person2
Creator Ana Fernandez Montraveta
Irene Castellón
Glòria Vázquez
Description The original SenSem Spanish Corpus includes syntactic and semantic annotations for a number of Spanish texts from the press domain developed by the GRIAL group (Grup de recerca consolidat de la Generalitat de Catalunya). The corpus contains one million words 300,000 of which were manually annotated at the syntactic and semantic level with syntagmatic categories, syntactic functions and semantic roles.
Language Spanish
Rights GPL
Source META-SHARE
Title GrAF version of the SenSem Spanish Corpus
Type Corpus
Contact Point Metashare/d7fbbd7c492e11e2a3fe0050569b00005c2192e3cce64d089f07447add8b4b2f#contact Person
Description Texts in the IT domain of the Danish DK-CLARIN LSP corpus come from Libris, Open Office, Aktuel Naturvidenskab. All texts are in XML TEIP5 format (TEIP5DKCLARIN-format), with tokenisation, pos-tagging, lemmatisation and termhood annotation placed text extermally in separate spangroups.
Language Danish
Rights CLARIN_ACA-NC
Source META-SHARE
Title DK-CLARIN LSP corpus - It domain
Type Corpus
Contact Point Metashare/72aee14ea37611e3960f001dd8b71c195970a482d03a42c3b9536622daaed5e3#contact Person
Description EASTIN-CL Multilingual Ontology of Assistive Technology was created within the EASTIN-CL project aimed at applying language technologies to portal of assistive technologies http://www.eastin.eu to enhance it and it more accessible for people in different languages. Based on Multilingual Ontology a query tool was built allowing users of the portal to type the lookup words which are then mapped to assistive device product classes. The terminology resource was created by first selecting base terminology in English, then having domain experts translate it into 6 other languages. The terminology resource is linked to ISO9999 classes. The current version 2 of it is aligned with ISO9999:2011 version of the classifier.
Language Italian
German
English
Rights CC-BY-SA
Source META-SHARE
Title EASTIN-CL Multilingual Ontology of Assistive Technology
Type Lexical Conceptual Resource
Contact Point Metashare/b7521676de7411e2b1e400259011f6eae05ca5a32e8742f3915120982df16bf3#contact Person
Description Technical domains Languages: Italian=>English Format: ASCII format with ISO 8859-1 character set Medium: QIC 150 MB Cartridge Tape Domain: Economics, 50,000 entries, canonical forms Technical bilingual Italian dictionaries with a morphological coding which can generate all full forms using a software engine written in C. Multi-word terms contain morphological coding for the head word.
Domaines techniques Langues : Italien=>Anglais Domaine: Economie, 50 000 entrées, formes canoniques Les dictionnaires techniques bilingues disposent d'une codification morphologique qui permet de générer toutes les formes fléchies grâce à un logiciel écrit en langage C. Les mots composés contiennent une codification morphologique sur la tête des mots.
Language English
Italian
Rights ELRA_VAR
ELRA_END_USER
Source META-SHARE
Title THAMUS Dictionnaires bilingues - Economie
THAMUS Bilingual dictionaries - Economics
Type Lexical Conceptual Resource
Contact Point Metashare/c6200aee92c211e28763000c291ecfc8de51a6dc181642d2943bb15705f33c33#contact Person2
Metashare/c6200aee92c211e28763000c291ecfc8de51a6dc181642d2943bb15705f33c33#contact Person
Description This is the LMF version of the Apertium bilingual dictionary for Occitan and Spanish languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as English-Catalan). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Language Spanish
Rights GPL
Source META-SHARE
Title Occitan-Spanish LMF Apertium Bilingual dictionary
Type Lexical Conceptual Resource
Contact Point Metashare/f05580e2de6b11e2b1e400259011f6ea21e29d2b658142ed9d6c69dfc5fce352#contact Person
Description * Entrées anglais-espagnol : Recherche scientifique & sciences mathématiques (906 entrées), géosciences (10 215), informatique, électronique & télécommunications (70 580), industrie (47 578), transports & maintenance (12 291), économie (145 572), sciences biologiques (38 989), communication & média (8 143), sciences chimiques & physiques (27 467). * Entrées allemand-anglais-espagnol-français : Environnement (36 658), santé (66 727), agriculture & alimentation (25 975), construction & travaux publics (8 429), droit & politique (56 578), sports & loisirs (17 312). * Deux lexiques spécialisés: Espagnol-anglais et allemand-anglais-français sans codes de domaine : électronique, télématique, droit, taxes, douanes, etc. (550 000 entrées). * Deux lexiques généraux: Allemand-anglais-espagnol-français et allemand-anglais-espagnol-français-italien-portugais. (83 000 entrées). Cette base de données terminologique contient, pour chaque domaine, l'indication de sous-domaines (de 2 sous-domaines pour la recherche scientifique à 39 pour les sports et loisirs). Chaque entrée comporte une définition, une unité phraséologique, une abréviation, une information sur l'usage et des étiquettes grammaticales. Format: ASCII Support : disquette
* Entries for English-Spanish: Scientific research & mathematical sciences (906 entries), Geosciences (10,215), Computer science, electronics & telecommunications (70,580), Industry (47,578), Transport & Maintenance (12,291), Economy (145,572), Biological sciences (38,989), Communication & media (8,143), Chemical & physical sciences (27,467). * Entries for English-French-German-Spanish: Environment (36,658), Health (66,727), Agriculture & food (25,975), Construction & public works (8,429), Law & policy (56,578), Sports & Leisure (17,312) * Two specialized lexicons: Spanish-English and English-French-German without domain codes: electronics, telematics, law, taxes, customs, etc. (550,000 entries). * Two general lexicons: Spanish-English-French-German and Spanish-English-French-German-Portuguese-Italian (83,000 entries). This terminological database contains, for each domain, a sub-domain indication is given (from 2 sub-domains for Scientific research to 39 for Sports & leisure). Each entry consists of a definition, phraseological unit, abbreviation, usage information, grammatical labels. Format: ASCII Medium: floppy disk
Language English
Spanish
Rights ELRA_VAR
Source META-SHARE
Title VERBA Polytechnic and Plurilingual Terminological Database - N-AX Economics of Real Estate
Base de données terminologique polytechnique et plurilingue VERBA - N-AX Economie de biens
Type Lexical Conceptual Resource
Contact Point Metashare/f2a91f6ede7711e2b1e400259011f6ea25642b0b191342bf8b43c73d5e286671#contact Person
Description This lexicon is subdivided into five different subsets: L0072-01 Full lexicon L0072-02 Phonetic layer L0072-03 Morphological layer L0072-04 Syntactic layer L0072-05 Semantic layer PAROLE-SIMPLE-CLIPS is a four-level, general purpose lexicon that has been elaborated over three different projects. The kernel of the morphological and syntactic lexicons was built in the framework of the LE-PAROLE project. The linguistic model and the core of the semantic lexicon were elaborated in the LE-SIMPLE project, while the phonological level of description and the extension of the lexical coverage were performed in the context of the Italian project Corpora e Lessici dell'Italiano Parlato e Scritto (CLIPS). The PAROLE-SIMPLE-CLIPS Pisa Italian Lexicon comprises a total of 387,267 phonetic units, 53,044 morphological units (53,044 lemmas), 37,406 syntactic units (28,111 lemmas) and 28,346 semantic units (19,216 lemmas). It was encoded at the semantic level, in full accordance with the international standards set out in the PAROLE-SIMPLE model and based on EAGLES. Syntactic and semantic encoding were performed jointly with Thamus (Consortium for Multilingual Documentary Engineering), which is responsible for 25,000 extra entries (to be released soon). PAROLE-SIMPLE-CLIPS offers therefore the advantage of being compatible with the other eleven PAROLE-SIMPLE lexicons that were built for European languages and that share a common theoretical model, representation language and building methodology. A PAROLE-SIMPLE-CLIPS entry gathers together all the phonological, morphological and inherent syntactic and semantic properties of a headword. Its subcategorization pattern is (or are) described in terms of optionality, syntactic function, syntagmatic realization as well as morpho-syntactic, syntactic and lexical properties of each slot filler. At the semantic level, the theoretical approach adopted by the SIMPLE model is essentially grounded on a revisited version of some fundamental aspects of the Generative Lexicon. A SIMPLE-CLIPS semantic unit is richly endowed with a wide range of fine-grained, structured information, most relevant for NLP applications. First among them, the ontological typing: the lexicon is in fact structured in terms of a multidimensional type system based on both hierarchical and non-hierarchical conceptual relations, taking into account the principle of orthogonal inheritance. Other relevant information types in a word entry are its domain of use; type of denoted event; synonymy and morphological derivation relations; membership in a class of regular polysemy as well as any relevant distinctive semantic features. Particularly outstanding is the information encoded in the Extended Qualia Structure (a set of 60 semantic relations that allow modelling both the different meaning dimensions of a word sense and its relationships to other lexical units) and the Predicative Representation which describes the semantic scenario the word sense considered is involved in and characterizes its participants in terms of thematic roles and semantic constraints. In a word’s description, lexical information is interrelated across the four description levels. Syntactic and semantic information, in particular, is related to each other through the projection of the predicate-argument structure onto its syntactic realization(s). References : Ruimy N., Corazzari O., Gola E., Spanu A., Calzolari N., Zampolli A. 2003. The PAROLE model and the Italian Syntactic lexicon. In A. Zampolli, N. Calzolari, L. Cignoni, (eds.), Computational Linguistics in Pisa - Linguistica Computazionale a Pisa. Linguistica Computazionale, Special Issue, XVIII-XIX, (2003). Pisa-Roma, IEPI. Tomo II, 793-820. Lenci A., Busa F., Ruimy N., Gola E., Monachini M., Calzolari N., Zampolli A. et al., 2000. SIMPLE Linguistic Specifications, SIMPLE LE4-8346 EC Project, Deliverable D2.1 & D2.2, WP02, Final version, March 2000, ILC and University of Pisa, 404 pp. (http://www.ub.es/gilcub/SIMPLE/simple.html#Specifications). Ruimy N., Monachini M., Gola E., Calzolari N., Del Fiorentino M.C., Ulivieri M., Rossi S. 2003. A computational semantic lexicon of Italian: SIMPLE. In A. Zampolli, N. Calzolari, L. Cignoni, (eds.), Computational Linguistics in Pisa - Linguistica Computazionale a Pisa. Linguistica Computazionale, Special Issue, XVIII-XIX, (2003). Pisa-Roma, IEPI. Tomo II, 821-864. Ruimy N., Monachini M., Distante R., Guazzini E., Molino S., Ulivieri M., Calzolari N., Zampolli A. 2002. CLIPS, A Multi-level Italian Computational Lexicon: a Glimpse to Data. LREC 2002. Las Palmas de Gran Canaria, Spain 29th, 30th & 31 May 2002. Proceedings, Volume III, Paris, The European Languages Resources Association (ELRA). 792-799.
Ce lexique est divisé en cinq sous-ensembles : L0072-01 Lexique complet L0072-02 Niveau phonétique L0072-03 Niveau morphologique L0072-04 Niveau syntaxique L0072-05 Niveau sémantique PAROLE-SIMPLE-CLIPS est un lexique générique à quatre niveaux qui a été élaboré au cours de trois projets différents. Le noyau des lexiques morphologique et syntaxique a été realisé dans le cadre du projet LE-PAROLE. Le modèle linguistique et le noyau du lexique sémantique ont été élaborés dans le projet LE-SIMPLE, tandis que le niveau phonologique de description et l’extension de la couverture lexicale ont été réalisés dans le contexte du projet italien Corpora e Lessici dell'Italiano Parlato e Scritto (CLIPS). Le lexique italien PAROLE-SIMPLE-CLIPS de Pise comprend un total de 387 267 unités phonétiques, 53 044 unités morphologiques (53 044 lemmes), 37 406 unités syntaxiques (28 111 lemmes) et 28 346 unités sémantiques (19 216 lemmes). Il a été codé au niveau sémantique, en respectant entièrement les standards internationaux fixés dans le modèle PAROLE-SIMPLE et basés sur EAGLES. Le codage syntaxique et sémantique ont été réalisés conjointement avec Thamus (Consortium pour l’ingénierie documentaire multilingue), qui est l’auteur de 25 000 entrées (à paraître). Ainsi, PAROLE-SIMPLE-CLIPS offre l’avantage d’être compatible avec les onze autres lexiques PAROLE-SIMPLE qui ont été construits pour les langues européennes et qui partagent un modèle théorique commun, un langage de représentation et une méthodologie de construction. Une entrée de type PAROLE-SIMPLE-CLIPS regroupe toutes les propriétés phonologiques, morphologiques et inhérentes à la syntaxe et à la sémantique d’un mot-tête (« headword »). Son modèle de sous-catégorisation est décrit en termes d’optionalité, de fonction syntaxique, de réalisation syntagmatique, ainsi qu’en termes de propriétés morpho-syntaxiques, syntaxiques et lexicales de chaque catégorie fonctionnelle (« slot-filler »). Au niveau sémantique, l’approche théorique adoptée par le modèle SIMPLE est essentiellement basée sur une version revisitée de quelques aspects fondamentaux du Lexique Génératif. Une unité sémantique SIMPLE-CLIPS est richement doté d’une grande variété d’informations fines et structurées, des plus importantes pour les applications en TAL. En tête de ces informations, la typologie ontologique : le lexique est en fait structuré en termes de systèmes de types multidimensionnels basé sur des relations conceptuelles hiérarchiques et non hiérarchiques, prenant en compte le principe d’héritage orthogonal. D’autres types d’information intéressants dans une entrée de mot sont son domaine d’usage, le type d’événement indiqué, la synonymie et les relations de dérivation morphologique, affectation à une classe de polysémie régulière, ainsi qu’à des traits sémantiques distinctifs. Une information particulièrement intéressante est l’information codée dans la Structure de Qualia étendue (un ensemble de 60 relations sémantiques qui permettente de modéliser à la fois les différentes dimensions de signification du sens d’un mot et ses relations avec les autres unités lexicales) et la Représentation prédicative qui décrit le scénario sémantique dans lequel est impliqué le sens du mot considéré et qui caractérise les participants en termes de rôles thématiques et de contraintes sémantiques. Dans une description de mot, l’information lexicale est étroitement liée entre les quatre nivaux de description. Les informations syntaxique et sémantique, en particulier, sont reliées entre elles grâce à la projection de la structure de l’argument-prédicat sur la ou ses réalisations syntaxiques. Références : Ruimy N., Corazzari O., Gola E., Spanu A., Calzolari N., Zampolli A. 2003. The PAROLE model and the Italian Syntactic lexicon. In A. Zampolli, N. Calzolari, L. Cignoni, (eds.), Computational Linguistics in Pisa - Linguistica Computazionale a Pisa. Linguistica Computazionale, Special Issue, XVIII-XIX, (2003). Pisa-Roma, IEPI. Tomo II, 793-820. Lenci A., Busa F., Ruimy N., Gola E., Monachini M., Calzolari N., Zampolli A. et al., 2000. SIMPLE Linguistic Specifications, SIMPLE LE4-8346 EC Project, Deliverable D2.1 & D2.2, WP02, Final version, March 2000, ILC et Université de Pisa, 404 pp. (http://www.ub.es/gilcub/SIMPLE/simple.html#Specifications). Ruimy N., Monachini M., Gola E., Calzolari N., Del Fiorentino M.C., Ulivieri M., Rossi S. 2003. A computational semantic lexicon of Italian: SIMPLE. In A. Zampolli, N. Calzolari, L. Cignoni, (eds.), Computational Linguistics in Pisa - Linguistica Computazionale a Pisa. Linguistica Computazionale, Special Issue, XVIII-XIX, (2003). Pisa-Roma, IEPI. Tomo II, 821-864. Ruimy N., Monachini M., Distante R., Guazzini E., Molino S., Ulivieri M., Calzolari N., Zampolli A. 2002. CLIPS, A Multi-level Italian Computational Lexicon: a Glimpse to Data. LREC 2002, Las Palmas de Gran Canaria, Espagne 29, 30 & 31 mai 2002. Proceedings, Volume III, Paris, The European Languages Resources Association (ELRA). 792-799.
Language Italian
Rights ELRA_VAR
ELRA_END_USER
Source META-SHARE
Title Lexique italien PAROLE-SIMPLE-CLIPS de Pise – Niveau morphologique
PAROLE-SIMPLE-CLIPS PISA Italian Lexicon – Morphological layer
Type Lexical Conceptual Resource