Index

Creator Universitat Pompeu Fabra (UPF)
Description Freeling-based chunker parser. Languages: English, Catalan, Spanish, Asturian and Galician. Input: Plain text Output: Freeling output format, XML, XML CQP ready. Input example: http://ws02.iula.upf.edu/panacea/examples/ws/freeling_parsed/freeling_parsed.input.example.txt Output example: http://ws02.iula.upf.edu/panacea/examples/ws/freeling_parsed/freeling_parsed.output.example.txt Output XML example: http://ws02.iula.upf.edu/panacea/examples/ws/freeling_parsed/freeling_parsed.output.xml Output XML CQP example: http://ws02.iula.upf.edu/panacea/examples/ws/freeling_parsed/freeling_parsed.output.cqp.xmlThis WS performs a FreeLing-based chunker parser (v 3.0). The WS requires a plain text input. The possible outputs formats are FreeLing , XML, and XML CQP ready. The languages supported are English, Catalan, Spanish, Asturian and Galician.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Corpus Processing, Syntactic Tagging,
Title Service - FreeLing Chunker parser Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Jimmy O'Regan
Description This is the LMF version of the Apertium bilingual dictionary for Spanish and Aragonese languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as Spanish-Aragonese). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Rights This resource is licensed under a GNU General Public License version 3.0 (http://www.gnu.org/licenses/gpl.html). The availability status of the resource is: 'available-unrestrictedUse'.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject 'language resources', 'lexical conceptual resource', 'bilingual lexicon'
Title Spanish-Aragonese LMF Apertium Bilingual dictionary
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description This is the tagging workflow for crawled data using Freeling. Freeling is run using the \"keeptags\" option to remove boilerplate and to keep paragraph tags info from the input data. The output is converted to the Travelling Object format (TO1 xces). This workflow uses the download_url processor for users who want to download the output files during the workfow execution and want the output files to follow the name of the input file. Example: 1234.xml > 1234.tag.xml. This workflow can be found without the download_url processor. Input files are read from a local directory.. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/35/versions/1/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, directory, folder, example, basicxces, freeling,
Title Freeling tagging for crawled data with input upload and output download
Creator Prompsit Language Engineering, S.L
Description This is the LMF version of the Apertium bilingual dictionary for Spanish Galician languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as English-Catalan). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Rights This resource is licensed under a GNU General Public License version 3.0 (http://www.gnu.org/licenses/gpl.html). The availability status of the resource is: 'available-unrestrictedUse'.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject 'language resources', 'lexical conceptual resource', 'bilingual lexicon'
Title Spanish-Galician LMF Apertium Bilingual dictionary
Creator Universitat Pompeu Fabra (UPF)
Description Processor to extract desired data by columns. Based on Linux awk. Columns: indicate the columns number you desire separated by commas. Input: Raw data. Default column separator is blank space or tabs. You can optionally specify the input and output separators. Example: Columns: 4,2 Input: http://ws02.iula.upf.edu/panacea/examples/ws/columns_selector/input_example_1.txt or http://ws02.iula.upf.edu/panacea/examples/ws/columns_selector/input_example_2.txt Output example: http://ws02.iula.upf.edu/panacea/examples/ws/columns_selector/output_example_1.txt This WS allows extracting a column from a tabular file input text. It is useful to work with CoNLL or FreeLing annotated corpora. Language independent WS.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Format Conversion,
Title Service - Columns selector Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description <p>This workflow uses FreeLing annotated data to classify the given list of nouns with the different available noun classifiers for Enlgish (7 classes). The LMF ouputs of each classifier are merged into a single LMF lexicon containing information for all classes.</p>. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/84/versions/2/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, merge, lexical acquisition, noun classification, english,
Title Classification of nouns in PoS tagged data for English and 7 available classes
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description <p>This workflow uses FreeLing annotated data to classify the given list of nouns with the different available noun classifiers for Spanish (9 classes). The LMF ouputs of each classifier are merged into a single LMF lexicon containing information for all classes.</p>. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/85/versions/2/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, merge, lexical acquisition, noun classification,
Title Classification of nouns in PoS tagged data for Spanish and 9 available classes
Creator Universitat Pompeu Fabra (UPF)
Description This WS calculates the probability of seeing a linguistic cue given a lexical class (P(cue|class) value). This probability is computed given the occurrences of cues in a corpus (codified in the signatures file) and the information of belonging or not belonging of these words to different classes (codified in indicators file). The probability is computed for each studied cue in the signatures file and for each class in the indicators file.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Statistics Analysis,
Title Service - P clue/ lexical class computer Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description Given a PoS tagged corpus, index it and get a list with all verbs with their frequency.. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/83/versions/1/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, cqp,
Title Get all verbs in PoS tagged corpus
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description <p>Example workflow to align sentences from a PDF documents (input consists of 2 lists of urls: lang A pdfs and lang B pdfs).</p>. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/21/versions/1/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, hunalign, panacea, sentence alignment, pdf,
Title PDF sentence alignment
Creator Hector Alòs i Font
Description This is the LMF version of the Apertium bilingual dictionary for Esperanto and Spanish languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as Esperanto-Spanish). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Rights This resource is licensed under a GNU General Public License version 3.0 (http://www.gnu.org/licenses/gpl.html). The availability status of the resource is: 'available-unrestrictedUse'.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject 'language resources', 'lexical conceptual resource', 'bilingual lexicon'
Title Esperanto-Spanish LMF Apertium Bilingual dictionary
Creator Universitat Politècnica de Catalunya. Research Group on Natural Language Processing
Description This is the stand-off GrAF version of Spanish portions of the Wikipedia (based on a 2006 dump). This Wikipedia Spanish Corpus contains 257019 articles that contain about 150,1 million words in raw text format. It has been cleaned by erasing disambiguation pages, removing some XML tags and homogenizing lists ending tag. Then, the corpus has been processed for adding structural tagging (head, paragraph, sentence, list, etc.) and morphosyntactic information.
Rights This resource is licensed under a GNU Free Documentation License 1.3 (http://www.gnu.org/copyleft/fdl.html). The availability status of the resource is: 'available-unrestrictedUse'.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject 'language resources', 'monolingual corpus'
Title GrAF version of Spanish portions of Wikipedia Corpus
Creator Universitat Pompeu Fabra (UPF)
Description This WS deploys a FreeLing-based text tokenizer (v 3.0). The WS splits a file in plain text format and UTF-8 encoded into units (tokens). The languages supported are Catalan, English, Galician, Italian, Portuguese, Russian, Spanish, Welsh, and Asturian.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Corpus Processing, Tokenization,
Title Service - FreeLing Tokenizer Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description This workflow uses UCAM SCF extractor to extract the SCF for a set of given verbs. The input corpus must be already parsed. The parameters used in this workflow assume that the corpus has been tagged (or follows the format) whit Spanish malt parser webservice as deployed by Panacea. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/80/versions/3/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, scf extraction, spanish,
Title Spanish SCF extractor from parsed corpus for a given list of verbs
Creator Universitat Pompeu Fabra (UPF)
Description This webservice performs traditional Naive Bayes classification of instances given in a weka file. It outputs the predicted classification for each instance and some statistics about the performance of the classification. The parameters needed as input can be learnt using estimate_bayesian_parameters webservice.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Lexicon/Terminology Extraction,
Title Service - Naive Bayes classifier Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra (UPF)
Description A WS to convert MS Word documents to plain text format. Language independent WS.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Format Conversion,
Title Service - MS Word to text converter Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra (UPF)
Description This WS will scramble the lines in a parallel text corpus keeping the alignment. The goal is to make it difficult to reproduce the original text. The input size limit is 100 MB. Language independent WS.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Alignment,
Title Service - Linescrambler parallel Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universitat Pompeu Fabra (UPF)
Description A morphosintatic tagger using a treetagger instance. Obligatory Inputs: -Plain text (or txt file) -Language (lang) es: Spanish ca: Catalan en: English Optional Inputs -Encoding UTF-8 (default) or ISO-8859-1 -keeptags If your text have a sgml tags and you will perserve it, mark this option -Output_form: treetagger (default) or iulact Input sample: -plain text http://ws02.iula.upf.edu/panacea/examples/ws/iula_tagger/plain-text.txtOutput sample: -iulact¹ http://ws02.iula.upf.edu/panacea/examples/ws/iula_tagger/outputIulaCT.txt -treetagger² http://ws02.iula.upf.edu/panacea/examples/ws/iula_tagger/outputTagger.txt Note: ¹iulact value in output_from ²treetagger value in output_fromThis WS is a morphosyntatic tagger. The disambiguation process is done by a TreeTagger instance trained by the IULA. The input is plain text in Catalan or Spanish. The output allows optional formats and optional encoding.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Morphosyntactic Tagging, Corpus Processing, Stemming/Lemmatization,
Title Service - IULA tagger Web Service
Type http://purl.org/net/def/metashare#toolService
Services
Creator Universidá d'Uviéu. Equipu d'Investigación Eslema
Description This is the LMF version of the Apertium bilingual dictionary for Spanish and Asturian languages. Bilingual LMF dictionaries were generated from Apertium bilingual dix files. For each Apertium bilingual correspondence, the corresponding source and target monolingual entries (LexicalEntry) were generated in addition to the bilingual correspondence (SenseAxis) element. Apertium is a free/open-source machine translation platform, initially aimed at related-language pairs but recently expanded to deal with more divergent language pairs (such as English-Catalan). The platform provides: a language-independent machine translation engine; tools to manage the linguistic data necessary to build a machine translation system for a given language pair and linguistic data for a growing number of language pairs.
Rights This resource is licensed under a GNU General Public License version 3.0 (http://www.gnu.org/licenses/gpl.html). The availability status of the resource is: 'available-unrestrictedUse'.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject 'language resources', 'lexical conceptual resource', 'bilingual lexicon'
Title Spanish-Asturian LMF Apertium Bilingual dictionary
Creator Universitat Pompeu Fabra. Institut Universitari de Lingüística Aplicada (IULA)
Description <p>This workflow is used to show how to handle lists when there's a processor which has different inputs with different depth lists.</p>. An image preview of the workflow can be found at: http://myexperiment.elda.org/workflows/3/versions/2/previews/full
Rights by-sa
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP workflow, Taverna 2, example, lists,
Title List example 01
Creator Universitat Pompeu Fabra (UPF)
Description This WS identifies human nouns in a part of speech tagged text (with FreeLing Morphosyntactic tagger V 3.0 WS). The classification is performed with a pre-trained Decision Tree. The ouptut is a LMF file with the classifier prediction for each noun. ou can choose to have this prediction as: - \"scored\": each noun gets a score of being or not being a member of the class (bigger than 0 means class member, smaller, non member of the class) - \"filtered\": the nouns are filtered according to their score. If the score is positive and over a determined threshold the noun is considered to be a member of the class. If it is negative and under another threshold, it is considered to be a non-member of the class. The other cases are tagged as \"unknown\", since the classifier did not give enough confidence to their classification. The used thresholds are pre-set according to some experiments, if you want to use your own thresholds, you should get the \"scored\" output and use the \"filter\" webservice to filter it with your thresholds. The languages supported are Spanish and English.
Source IULA UPF Centre de Competencia CLARIN (via CLARIN VLO)
Subject NLP service, Lexicon/Terminology Extraction,
Title Service - Human nouns classifier Web Service
Type http://purl.org/net/def/metashare#toolService
Services