Index

Creator Datahub/apertium-rdf-fr-es#N6b21c2082318420c821155761a318fd3
Description RDF version of the Apertium bilingual dictionary FR-ES. The original dataset (in LMF) comes from http://hdl.handle.net/10230/17090. The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF FR-ES
Creator Datahub/apertium-rdf-oc-ca#N4bd161bc6d974053a3ffa84be2819e4c
Description RDF version of the Apertium bilingual dictionary OC-CA. The original dataset (in LMF) comes from http://hdl.handle.net/10230/17107 . The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF OC-CA
Creator Datahub/apertium-rdf-oc-es#N644cc27b02574af49889778652b16e83
Description RDF version of the Apertium bilingual dictionary OC-ES. The original dataset (in LMF) comes from <http://hdl.handle.net/10230/17124> . The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF OC-ES
Creator Datahub/apertium-rdf-pt-ca#N07be4e0c3ce74a99a76117f7ce0c67ec
Description RDF version of the Apertium bilingual dictionary PT-CA. The original dataset (in LMF) comes from http://hdl.handle.net/10230/17120 . The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF PT-CA
Creator Datahub/apertium-rdf-pt-gl#N44d656dd11954b35bb595e37a7d14b5d
Description RDF version of the Apertium bilingual dictionary PT-GL. The original dataset (in LMF) comes from http://hdl.handle.net/10230/17115 . The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF PT-GL
Creator Datahub/apertium-rdf#Nd3cd20b01b3049a3a27d20773b75f63d
Description This dataset groups all the Apertium RDF bilingual dictionaries (see the whole list at http://linghub.lider-project.eu/datahub?q=apertium+rdf&organization=oeg-upm). The dictionaries were converted into RDF from their LMF version, which can be found in Meta-Share (http://metashare.upf.edu/). The RDF version of the lexica is modelled using lemon (http://lemon-model.net/) and the translation module (http://purl.org/net/translation). The Apertium RDF version has been generated by OEG (Universidad Politécnica de Madrid) jointly with IULA (Universitat Pompeu Fabra). More about the Apertium RDF dictionaries at http://linguistic.linkeddata.es/apertium/
Source DataHub
Title Apertium RDF
Creator Datahub/asit#Naed6ee713bbe4e94ae1cdb0c2937fb79
Description The Atlante Sintattico d'Italia, Syntactic Atlas of Italy (ASIt) enterprise builds on a long standing tradition of collecting and analysing linguistic corpora, which has originated different efforts and projects over the years. ASIt accounts for minimally different variants within a sample of closely related languages, thus it does not need a thorough part of speech (POS) disambiguation, since the \"trivial\" identification of basic POS (e.g. Nouns vs Verbs) is not enough to capture cross-linguistic differences between closely related languages. Secondly, the linguistic variants cannot be reduced to lexical distinctions only, i.e. syntactic differences are in general unpredictable on the basis of the properties of single lexical items. A specific tag set designed to capture sentence-level phenomena without taking into consideration POS tags is needed. As a consequence, while other tag sets are designed to carry out a gross linguistic analysis of a vast corpus, the ASIt tag set aims to capture fine-grained grammatical differences by comparing various dialectal translations of the same sentence. Moreover, in order to pin down these subtle asymmetries, the linguistic analysis must be carried out manually. To explain why the needs for ASIt are so special we have to take into consideration two different aspects: the nature of Italian dialects, and the kind of linguistic theory ASIt aims to interact with. The Italian dialectal area presents a kind of variation that involves parametric choices affecting many general aspects of syntax, morphology, and phonology. The kind of information we want to gather involves not only the presence of a certain element, but also the absence of an element; an element can be omitted only in some constructions and in conjunction with specific characteristics of the language. For this reason, ASIt proposed the creation of a specific set of tags starting from a universal core shared by all languages (on the basis of the work done by DynaSAND), and subsequently developing a language-specific periphery which is compatible with other projects. Dialectal data stored in the ASIt were gathered during a twenty-year-long survey investigating the distribution of several grammatical phenomena across the dialects of Italy. These data and information were collected by means of questionnaires formed by sets of Italian sentences: dialectal speakers were asked to translate them into their dialects and write their translations in the questionnaire; therefore, each questionnaire is associated with many parallel dialectal translations. At present, there are eight different questionnaires written in Italian and almost 500 questionnaires, corresponding to the eight Italian questionnaires, written in more than 240 different dialects, for a total of more than 54,000 sentences and more than 40,000 tags stored in the data resource managed by the ASIt digital library system.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title Atlante Sintattico d'Italia (ASIt)
Creator Datahub/babelnet#N866ae584266049c08155783eaf271ca8
Description BabelNet is both a multilingual encyclopedic dictionary, with lexicographic and encyclopedic coverage of terms, and an ontology which connects concepts and named entities in a very large network of semantic relations, made up of 13,801,844 millions of nodes, called Babel synsets. Each Babel synset represents a given meaning and contains all the synonyms which in different languages express that meaning. The BabelNet resource is made available under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 Unported license. The different resources from which BabelNet originates are made available under different licenses, as follows: WordNet: http://wordnet.princeton.edu/wordnet/license/ , Wikipedia: Creative Commons Attribution-ShareAlike 3.0 Unported license, OmegaWiki: Creative Commons Attribution 2.5 Generic license or higher, and Open Multilingual WordNet: different licenses according to the language, as indicated at http://compling.hss.ntu.edu.sg/omw/ . When applicable, specific license rights are specified through the property dcterms:license. Please make sure of using data in compliance with their respective licenses.
Source DataHub
Title BabelNet
Creator Datahub/brown-corpus-in-rdf-nif#N9975bc86baa043718518170e40b547b3
Description RDF version of the Brown Corpus (W. N. Francis, H. Kucera; Brown University; 1979). 1,014,312 words in 500 documents, taken from newspapers texts on diverse topics, non-fiction and fiction books as well as government documents. Original corpus contains manually annotated sentence and token boundaries as well as word class annotations(such as POS, inflectional morphemes, such as noun plural, verb tense and adjective comparison and special tags for foreign words and proper nouns). Converted corpus contains complete texts reconstructed from TEI/XML version of the Brown corpus. Word classes where linked via OLiA to ontological categories for aggregated querying.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title Brown Corpus in RDF/NIF
Creator Datahub/chat-game-corpus#N6617551eedae4a93b77ce1cd78d9778c
Description A corpus resulting from an object arrangement game using a computer-mediated setting.
Rights http://www.opendefinition.org/licenses/odc-by
Source DataHub
Title Chat Game corpus
Creator Datahub/clean-energy-data-reegle#N50381869686f4f0b9f289ce5ef4e9508
Description Comprehensive set of linked clean energy data including: * policy and regulatory country profiles, * key stakeholders (organisation profiles), * project outcome documents and a * thesaurus (SKOS format) on renewables, energy efficiency and climate change for public re-use.
Rights http://reference.data.gov.uk/id/open-government-licence
Source DataHub
Title Linked Clean Energy Data (reegle.info)
Creator Datahub/core#Nbc9465d70d9c4472a54babc5395b2c41
Description The CORE dataset contains information about similarities between scientific papers stored across Open Access repositories. The similarities are calculated using Natural Language Processing techniques based on the full-text. The similarities are provided only for research articles with an accessible and machine readable full-text. More information about the data structure can be found at:http://core-project-local.kmi.open.ac.uk/data-description. #### RDF Statistics At the moment we expose more than 92 million RDF triples describing similarities calculated on a set of more than 400k full-text articles harvested from over 230 Open Access repositories. #### Links The data about the similarities are represented using the MuSIM ontology (http://kakapo.dcs.qmul.ac.uk/ontology/musim/0.2/musim.html) BIBO ontologies (http://bibliontology.com/) with links to the OAI (RKBExplorer) repository available in the Linked Data cloud.
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title CORE - Semantic Similarity of Open Access publications
Creator Datahub/cornetto#N665ac76ccf314bc184aee4cc1b1eece9
Description [Dutch lexical database](http://www2.let.vu.nl/oz/cltl/cornetto/), similar to WordNet but with more semantic relations. Links to package:vu-wordnet and package:w3c-wordnet . When this dataset is used for research purposes, please cite: Vossen, P., Maks, I., Segers, R., van der Vliet, H.: Integrating lexical units, synsets and ontology in the Cornetto database. In: (ELRA), E.L.R.A. (ed.) Proceedings of the Sixth International Language Resources and Evaluation (LREC’08) (2008)
Source DataHub
Title Cornetto1.2
Creator Datahub/dbnary#Nf82de97a0daa4f518eac989b52108162
Description Extracts of wiktionary data for several languages, structured as an RDF graph, based mainly on the LEMON model. English, Finnish, French, German, Greek, Italian, Japanese, Portuguese, Russian and Turkish.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title dbnary
Creator Datahub/dbpedia-abstract-corpus#N5b719f1e67ea4f4683b230acc01c8ad9
Description This corpus contains a conversion of Wikipedia abstracts in six languages (dutch, english, french, german, italian and spanish) into the I used the NLP Interchange Format (NIF). The corpus contains the abstract texts, as well as the position, surface form and linked article of all links in the text. As such, it contains entity mentions manually disambiguated to Wikipedia/DBpedia resources by native speakers, which predestines it for NER training and evaluation. Furthermore, the abstracts represent a special form of text that lends itself to be used for more sophisticated tasks, like open relation extraction. Their encyclopedic style, following Wikipedia guidelines on opening paragraphs adds further interesting properties. The first sentence puts the article in broader context. Most anaphers will refer to the original topic of the text, making them easier to resolve. Finally, should the same string occur in different meanings, Wikipedia guidelines suggest that the new meaning should again be linked for disambiguation. In short: The type of text is highly interesting. Acknowledgments: The conversion of this corpus was supported by the [FREME H2020 project](http://www.freme-project.eu/).
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title DBpedia abstract corpus
Creator Datahub/dbpedia-fr#N9364835b5a7c4e62817215a464e4389f
Description DBpedia in French dataset. Part of the DBpedia internationalisation effort. Data are extracted here from French speaking pages of wikipedia.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia in French
Creator Datahub/dbpedia-it#N86a2d373191647adad247a67687fb252
Description DBpedia is a \"community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia, and to link other data sets on the Web to Wikipedia data. We hope this will make it easier for the amazing amount of information in Wikipedia to be used in new and interesting ways, and that it might inspire new mechanisms for navigating, linking and improving the encyclopaedia itself.\"
Source DataHub
Title DBpedia in Italian
Creator Datahub/dbpedia-live#N29a21abc36424dd395785d62563b3edc
Description DBpedia.org is a community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia and to link other datasets on the Web to Wikipedia data. The DBpedia knowledge base currently describes more than 3.4 million things, out of which 1.5 million are classified in a consistent Ontology, including 312,000 persons, 413,000 places, 94,000 music albums, 49,000 films, 15,000 video games, 140,000 organizations, 146,000 species and 4,600 diseases. The DBpedia data set features labels and abstracts for these 3.2 million things in up to 92 different languages; 841,000 links to images and 5,081,000 links to external web pages; 9,393,000 external links into other RDF datasets, 565,000 Wikipedia categories, and 75,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1 billion pieces of information (RDF triples) out of which 257 million were extracted from the English edition of Wikipedia and 766 million were extracted from other language editions.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia-Live
Creator Datahub/dbpedia-nl#N3fd5515acd1447ec8f295fa84ef40b5c
Description DBpedia is a \"community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia, and to link other data sets on the Web to Wikipedia data. We hope this will make it easier for the amazing amount of information in Wikipedia to be used in new and interesting ways, and that it might inspire new mechanisms for navigating, linking and improving the encyclopaedia itself.\"
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia in Dutch
Creator Datahub/dbpedia-pt#Nc1d72da3f200412194e990d5fe46e9ef
Description DBpedia Portuguese is being constructed by a working group to internationalize DBpedia to the [Lusosphere](http://en.wikipedia.org/wiki/Lusosphere \"what is lusosphere?\"). We are actively collaborating with the [DBpedia Internationalization Team](http://wiki.dbpedia.org/Internationalization) to include knowledge from the [Portuguese Language Wikipedia](http://pt.wikipedia.org) into DBpedia. As a first step, we have performed a preliminary extraction (available at http://pt.dbpedia.org) and are editing several mappings and labels to the DBpedia Ontology. The activity is organized via the mailing list dbpedia-portuguese@lists.sourceforge.net http://pt.dbpedia.org
Source DataHub
Title DBpedia in Portuguese