Index

Creator Datahub/dbnary#Nf82de97a0daa4f518eac989b52108162
Description Extracts of wiktionary data for several languages, structured as an RDF graph, based mainly on the LEMON model. English, Finnish, French, German, Greek, Italian, Japanese, Portuguese, Russian and Turkish.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title dbnary
Creator Datahub/dbpedia-abstract-corpus#N5b719f1e67ea4f4683b230acc01c8ad9
Description This corpus contains a conversion of Wikipedia abstracts in six languages (dutch, english, french, german, italian and spanish) into the I used the NLP Interchange Format (NIF). The corpus contains the abstract texts, as well as the position, surface form and linked article of all links in the text. As such, it contains entity mentions manually disambiguated to Wikipedia/DBpedia resources by native speakers, which predestines it for NER training and evaluation. Furthermore, the abstracts represent a special form of text that lends itself to be used for more sophisticated tasks, like open relation extraction. Their encyclopedic style, following Wikipedia guidelines on opening paragraphs adds further interesting properties. The first sentence puts the article in broader context. Most anaphers will refer to the original topic of the text, making them easier to resolve. Finally, should the same string occur in different meanings, Wikipedia guidelines suggest that the new meaning should again be linked for disambiguation. In short: The type of text is highly interesting. Acknowledgments: The conversion of this corpus was supported by the [FREME H2020 project](http://www.freme-project.eu/).
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title DBpedia abstract corpus
Creator Datahub/dbpedia-fr#N9364835b5a7c4e62817215a464e4389f
Description DBpedia in French dataset. Part of the DBpedia internationalisation effort. Data are extracted here from French speaking pages of wikipedia.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia in French
Creator Datahub/dbpedia-live#N29a21abc36424dd395785d62563b3edc
Description DBpedia.org is a community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia and to link other datasets on the Web to Wikipedia data. The DBpedia knowledge base currently describes more than 3.4 million things, out of which 1.5 million are classified in a consistent Ontology, including 312,000 persons, 413,000 places, 94,000 music albums, 49,000 films, 15,000 video games, 140,000 organizations, 146,000 species and 4,600 diseases. The DBpedia data set features labels and abstracts for these 3.2 million things in up to 92 different languages; 841,000 links to images and 5,081,000 links to external web pages; 9,393,000 external links into other RDF datasets, 565,000 Wikipedia categories, and 75,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1 billion pieces of information (RDF triples) out of which 257 million were extracted from the English edition of Wikipedia and 766 million were extracted from other language editions.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia-Live
Creator Datahub/dbpedia-nl#N3fd5515acd1447ec8f295fa84ef40b5c
Description DBpedia is a \"community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia, and to link other data sets on the Web to Wikipedia data. We hope this will make it easier for the amazing amount of information in Wikipedia to be used in new and interesting ways, and that it might inspire new mechanisms for navigating, linking and improving the encyclopaedia itself.\"
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia in Dutch
Creator Datahub/dbpedia#Nf7838f5dc3c24b88b435398442a0b465
Description ### Description From the front page: > DBpedia.org is a community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia and to link other datasets on the Web to Wikipedia data. > > The DBpedia knowledge base currently describes more than 3.64 million things, out of which 1.83 million are classified in a consistent Ontology, including 416,000 persons, 526,000 places, 106,000 music albums, 60,000 films, 17,500 video games, 169,000 organisations, 183,000 species and 5,400 diseases. The DBpedia data set features labels and abstracts for these 3.64 million things in up to 97 different languages; 2,724,000 links to images and 6,300,000 links to external web pages; 6,200,000 external links into other RDF datasets, 740,000 Wikipedia categories, and 2,900,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1.2 billion pieces of information (RDF triples) out of which 335 million were extracted from the English edition of Wikipedia and 865 million were extracted from other language editions.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title DBpedia
Creator Datahub/dbpedia-spotlight-nif-ner-corpus#Na70340f2fbf44ff28ce4e219af2cb7eb
Description Based on P. N. Mendes, M. Jakob, A. García-Silva, and C. Bizer. DBpedia Spotlight: shedding light on the web of documents. In Proc. of the 7th Int. Conf. on Semantic Systems, 2011. It contains 60 natural language sentences from ten different New York Times articles with overall 249 annotated DBpedia entities, i. e. the entities are not explicitely bound to mentions within the texts, which causes a certain lack of clarity. Therefore, we (in all conscience) retroactively have allocated the entities to their positions within the texts. The entities dbp:Markup_Language and dbp:PBC_CSKA_Moscow could not be linked in the texts, since there was also a more specific entity enlisted occupying their solely possible location, e. g. hypertext markup language has been annotated with dbp:HTML rather than dbp:Markup_language.
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title DBpedia Spotlight NIF NER Corpus
Description DBpedia Spotlight is a tool for annotating mentions of DBpedia resources in text, providing a solution for linking unstructured information sources to the Linked Open Data cloud through DBpedia. DBpedia Spotlight performs named entity extraction, including entity detection and Name Resolution (a.k.a. disambiguation). It can also be used for building your solution for Named Entity Recognition, amongst other information extraction tasks. The datasets you find here were produced by the DBpedia Spotlight team and can be reused in many natural language processing tools, as well as general Web applications that need to connect text to unique URIs from DBpedia.
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title DBpedia Spotlight
Creator Datahub/environmental-applications-reference-thesaurus#N5bb1ca2f37ec4e879042c0e2426c5cfe
Description The Environmental Applications Reference Thesaurus (EARTh) has been compiled and is maintained by the CNR-IIA-EKOLab to facilitate the indexing, retrieval, harmonizing and integration of human- and machine-readable environmental information from disparate sources, across the cultural and linguistic barriers. Ownership of such material always remains with the CNR-IIA-EKOLab. EARTh has been firstly made available as linked data as an activity within the European Project NatureSDIPlus (ECP-2007-GEO-317007). It is currently maintained in the context of LusTRE, a framework under development within the EU project eENVplus (CIP-ICT-PSP grant No. 325232) that aims at combining existing thesauri to support the management of environmental resources. LusTRE considers the heterogeneity in scopes and levels of abstraction of environmental thesauri as an asset when managing environmental data, it exploits linked data best practices SKOS (Simple Knowledge Organization System) and RDF (Resource Description Framework) in order to provide a multi-thesauri solution for INSPIRE data themes related to the environment.
Rights http://creativecommons.org/licenses/by-nc/2.0/
Source DataHub
Title EARTh
Creator Datahub/eurosentiment#Ne49e1534eb9b4fa6846d42e3fcf3327d
Description Gabriela Vulcu, Raul Lario Monje, Mario Munoz, Paul Buitelaar and Carlos A. Iglesias (2014), Linked-Data based Domain-Specific Sentiment Lexicons, In: Proceedings of the 3rd Workshop on Linked Data in Linguistics (LDL-2014), Reykjavik, Iceland, May 2014 Resource Type: Lexicon Resource Name: EUROSENTIMENT domain-specific sentiment lexicons Size: 9160 Resource Production Status: Newly created-in progress Language(s): W Modality: Written Use of the Resource: Emotion Recognition/Generation Resource Availability: Freely Avalable Resource URL (if available): http://140.203.155.231:8080/eurosentiment/ Resource Description: Please see the submitted article. It describes this language resource and all other involved.
Rights http://www.opendefinition.org/licenses/gfdl
Source DataHub
Title EuroSentiment
Creator Datahub/fao-geopolitical-ontology#N520a3859e3d8418fbee243296fa13db6
Description The FAO geopolitical ontology provides a master reference for geopolitical information, as it manages names in multiple languages (English, French, Spanish, Arabic, Chinese, Russian and Italian); maps standard coding systems (UN, ISO, FAOSTAT, AGROVOC, DBPedia, etc); provides relations among territories (land borders, group membership, etc); and tracks historical changes. The ontology contains *number of triples : 22495 triples *links to other data sets: 195 links to DBPEDIA The Food and Agriculture Organization of the United Nations (FAO) leads international efforts to defeat hunger and serves as a knowledge network.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title FAO geopolitical ontology
Creator Datahub/fiesta#Nb6b2f4b3ba764141931f4f1a69617789
Description FiESTA (short for \"Format for extensive spatiotemporal annotations\") is a generic format for linguistic and behavioral annotations.
Rights http://www.opendefinition.org/licenses/cc-by-sa
Source DataHub
Title FiESTA
Creator Datahub/french-timebank#Nd85865e038964dd2a23c7b4ed6bd36cb
Description The French TimeBank consists of a set of 109 journalistic articles from 7 different sub-genres annotated according to the ISO-TimeML standard, adapted for the French language. Eventualities (events and states) and temporal expressions (dates, durations, frequencies, quantified intervals) are marked up with in-line annotation. The temporal relations that hold among these entities are also annotated, as are the aspectual and modal subordination relations between eventualities. The corpus is available under the Lesser General Public Licence for Linguistic Resources (LGPL-LR).
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title French TimeBank
Description This is the Galician EuroWordNet-Lemon lexicon. The lexicon was created from the Spanish Word-Net-LMF lexicon which is part of the Multilingual Central Repository (MCR http://adimen.si.ehu.es/web/MCR). The lexicon conforms to the 'lemon' specification. Gloss and rgloss relations between synsets are not included. For LexicalEntries and LexicalSenses, original ID's are encoded in dcterms:source and the URIs follow the pattern '../lemma-PoS'. For Synsets and Translations original IDs are used in the URIs (.../ID). Synset rdfs:labels were generated as follows: INSERT {?synset rdfs:label ?labels } WHERE {SELECT ?synset (GROUP_CONCAT(?label; separator = ' ; ') as ?labels) { ?sense lemon:reference ?synset; rdfs:label ?label . } GROUP BY ?synset } http://lodserver.iula.upf.edu/id/WordNetLemon/GL/
Rights http://www.opendefinition.org/licenses/cc-by
Source DataHub
Title Galician EuroWordNet-lemon lexicon (3.0)
Creator Datahub/gemeenschappelijke-thesaurus-audiovisuele-archieven#N5f86544c9f384fd8a38f8ea842f161dd
Description The Netherlands Institute for Sound and Vision <http://portal.beeldengeluid.nl/> is the Dutch archive for public broadcast television. They employ the GTAA, which is a Dutch acronym for Common Thesaurus [for] Audiovisual Archives, to index and disclose their audiovisaul documents. The GTAA closely follows the ISO-2788 standard for thesaurus structures. The thesaurus consists of several facets for describing TV programs: subjects; people mentioned; named entities (Corporation names, music bands etc); locations; genres; makers and presentators. The GTAA contains approximately 160.000 terms: ~3800 Subjects, ~97.000 Persons, ~27.000 Names, ~14.000 Locations, 113 Genres and ~18.000 Makers, and is continually updated as new concepts emerge on TV.
Rights http://www.opendefinition.org/licenses/odc-odbl
Source DataHub
Title Gemeenschappelijke Thesaurus Audiovisuele Archieven – Common Thesaurus Audiovisual Archives
Creator Datahub/gemet-annotated#Ndada529569ca4b958359f9509831cdd6
Description Details about how this dataset was built are described in the article: Are SKOS concept schemes ready for multilingual retrieval applications? — Diana Tanase and Epaminondas Kapetanios
Rights http://www.opendefinition.org/licenses/odc-odbl
Source DataHub
Title gemet-annotated