The Swedish N-grams 1770-1940 of the Newspaper and Periodical Corpus of the National Library of Finland

Instance of: Dataset
Description The National Library of Finland has digitized a large proportion of Finland’s Swedish newspapers, magazines, and periodicals published between 1770 and 1940. This resource contains sets of unigrams, bigrams and trigrams extracted from a corpus that has been compiled from the digitized newspapers by the University of Helsinki. The resource consists of plain UTF-8 encoded text files, each containing a list of n-grams that have been ordered by their frequencies from highest to lowest. Each line in a file consists of two or more fields separated by a whitespace character. The first field indicates the absolute frequency of a unique n-gram, and the remaining fields contain the tokens (strings of non-whitespace characters) of the n-gram itself. Uppercase letters have been retained as such and have not been converted into lowercase letters. Punctuation characters are treated as separate tokens except when they are part of an abbreviation (\"etc.\", \"mm.\"). The n-grams have been computed across sentence boundaries for each decade (from the 1770s to the 1940s) as well as for the entire corpus, with unigrams, bigrams and trigrams in separate files. Since the source material has been digitized by the means of optical character recognition (OCR), the resource also contains erroneous word forms and non-word strings of characters. Furthermore, due to the large time span of the original corpus, the resource contains several lexical items and spelling variants that have since become obsolete in standard Swedish. The resource will be updated in the future as improvements are being made to the source material. Referring to the Swedish N-gram Corpus If you use material from the Swedish N-gram Corpus and want to quote it, you may want to use the following information: Bibliographic references The Swedish N-gram Corpus, version 1 (SNC1). 2014. Distributed by the University of Helsinki on behalf of the FIN-CLARIN Consortium. URL: http://www.helsinki.fi/finclarin/snc1 Data from the SNC1 Our policy is to request that citations from the Swedish N-gram Corpus should include the corpus identifier and version number (a 4 letter code). A suitable way of crediting the SNC1 would be: \"N-grams from the Swedish N-gram Corpus, version 1, (SNC1) and the frequencies derived from it were obtained under the CC BY 4.0 license.\"
Language sv
Language Swedish
Rights CC-BY
See Also http://metashare.elda.org/repository/browse/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154/
Source META-SHARE
Title The Swedish N-grams 1770-1940 of the Newspaper and Periodical Corpus of the National Library of Finland
Type Resource Info
Type Corpus

Contact Point

Communication Info
Email [email protected]
Type Communication Info
Given Name User support
Surname FIN-CLARIN
Type Contact Person
Person
Person Info Type

Corpus Info

Corpus Text Ngram Info
Character Encoding Info
Character Encoding UTF-8
Type Character Encoding Info
Geographic Coverage Info
Geographic Coverage Finland
Type Geographic Coverage Info
Language Info
Language Swedish
Language sv
Language Name Swedish
Type Language Info
Linguality Info
Linguality Type Monolingual
Type Linguality Info
Media Type Text Ngram
Modality Info
Modality Type Written Language
Type Modality Info
Ngram Info
Base Item Word
Order 3 Int
Type Ngram Info
Size Info
Size 10558
Size Unit Mb
Type Size Info Type
Time Coverage Info
Time Coverage 1770-1949
Type Time Coverage Info
Type Corpus Text Ngram Info
Resource Type Corpus
Type Corpus Info

Distribution Info

Availability Available-unrestricted Use
Ipr Holder
Organization Info
Communication Info
Email [email protected]
Type Communication Info
Email [email protected]
Type Communication Info
Organization Name University of Helsinki
Organization Short Name UHEL
Type Organization Info Type
Type Actor
License
Attribution Text The Swedish N-gram Corpus, version 1 (SNC1). 2014. Distributed by the University of Helsinki on behalf of the FIN-CLARIN Consortium. URL: http://www.helsinki.fi/finclarin/snc1
Delivery Channel Downloadable
Distribution Rights Holder
Organization Info Metashare/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154#organization Info2
Type Actor
Licensor
Organization Info Metashare/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154#organization Info2
Type Actor
Permission
Action http://creativecommons.org/ns/Attribution
Duty Metashare/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154#permission
Type Duty
Permission
Restrictions Of Use
Same As https://creativecommons.org/licenses/by/4.0/
Type Licence Info
Type Distribution
Distribution Info

Identification Info

Description The National Library of Finland has digitized a large proportion of Finland’s Swedish newspapers, magazines, and periodicals published between 1770 and 1940. This resource contains sets of unigrams, bigrams and trigrams extracted from a corpus that has been compiled from the digitized newspapers by the University of Helsinki. The resource consists of plain UTF-8 encoded text files, each containing a list of n-grams that have been ordered by their frequencies from highest to lowest. Each line in a file consists of two or more fields separated by a whitespace character. The first field indicates the absolute frequency of a unique n-gram, and the remaining fields contain the tokens (strings of non-whitespace characters) of the n-gram itself. Uppercase letters have been retained as such and have not been converted into lowercase letters. Punctuation characters are treated as separate tokens except when they are part of an abbreviation (\"etc.\", \"mm.\"). The n-grams have been computed across sentence boundaries for each decade (from the 1770s to the 1940s) as well as for the entire corpus, with unigrams, bigrams and trigrams in separate files. Since the source material has been digitized by the means of optical character recognition (OCR), the resource also contains erroneous word forms and non-word strings of characters. Furthermore, due to the large time span of the original corpus, the resource contains several lexical items and spelling variants that have since become obsolete in standard Swedish. The resource will be updated in the future as improvements are being made to the source material. Referring to the Swedish N-gram Corpus If you use material from the Swedish N-gram Corpus and want to quote it, you may want to use the following information: Bibliographic references The Swedish N-gram Corpus, version 1 (SNC1). 2014. Distributed by the University of Helsinki on behalf of the FIN-CLARIN Consortium. URL: http://www.helsinki.fi/finclarin/snc1 Data from the SNC1 Our policy is to request that citations from the Swedish N-gram Corpus should include the corpus identifier and version number (a 4 letter code). A suitable way of crediting the SNC1 would be: \"N-grams from the Swedish N-gram Corpus, version 1, (SNC1) and the frequencies derived from it were obtained under the CC BY 4.0 license.\"
Distribution
Access URL http://urn.fi/urn:nbn:fi:lb-2014091903
Type Distribution
URL
Identifier http://urn.fi/urn:nbn:fi:lb-2014091902
Meta Share Id NOT_DEFINED_FOR_V2
Resource Short Name SNC1
Title The Swedish N-grams 1770-1940 of the Newspaper and Periodical Corpus of the National Library of Finland
Type Identification Info

Relation Info

Related Resource http://metashare.csc.fi/repository/browse/the-newspaper-and-periodical-corpus-of-the-national-library-of-finland/1fbbd932de3811e28c39005056be118e932a745cd63943beabebfe6cd27478fb/
Relation Type N-gramsOf
Type Relation Info

Usage Info

Actual Use Info
Actual Use Human Use
Type Actual Use Info
Use NLPSpecific Linguistic Research
Foreseen Use Info
Foreseen Use Human Use
Type Foreseen Use Info
Use NLPSpecific Linguistic Research
Type Usage Info

Metashare/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154#Header

Instance of: Catalog Record
Issued 2014-09-29T19:03:31Z Date
Primary Topic The Swedish N-grams 1770-1940 of the Newspaper and Periodical Corpus of the National Library of Finland
Set Spec corpus:textngram
corpus

Metashare/52380cb43ca111e48f80005056be118ea6e21787d4cf4f6999549fe57fc54154#metadata Info

Instance of: Catalog Record
Created 2014-09-15 Date
Language en
Language English
Language Name English
Modified 2014-09-15 Date
Primary Topic The Swedish N-grams 1770-1940 of the Newspaper and Periodical Corpus of the National Library of Finland
Type Metadata Info

Creator

Type Actor