Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach

Abstract: Nasa Yuwe is the language of the Nasa indigenous community in Colombia. It is currently threatened with extinction. In this regard, a range of computer science solutions have been developed to the teaching and revitalization of the language. One of the most suitable approaches is the construction of a Part-Of-Speech Tagging (POST), which encourages the analysis and advanced processing of the language. Nevertheless, for Nasa Yuwe no tagged corpus exists, neither is there a POS Tagger and no related works have been reported. This paper therefore concentrates on building a linguistic corpus tagged for the Nasa Yuwe language and generating the first tagging application for Nasa Yuwe. The main results and findings are 1) the process of building the Nasa Yuwe corpus, 2) the tagsets and tagged sentences, as well as the statistics associated with the corpus, 3) results of two experiments to evaluate several POS Taggers (a Random tagger, three versions of HSTAGger, a tagger based on the harmony search metaheuristic, and three versions of a memetic algorithm GBHS Tagger, based on Global-Best Harmony Search (GBHS), Hill Climbing and an explicit Tabu memory, which obtained the best results in contrast with the other methods considered over the Nasa Yuwe language corpus.

Saved in:
Bibliographic Details
Main Authors: Sierra Martínez,Luz Marina, Cobos,Carlos Alberto, Corrales Muñoz,Juan Carlos, Rojas Curieux,Tulio, Herrera-Viedma,Enrique, Peluffo-Ordóñez,Diego Hernán
Format: Digital revista
Language:English
Published: Instituto Politécnico Nacional, Centro de Investigación en Computación 2018
Online Access:http://www.scielo.org.mx/scielo.php?script=sci_arttext&pid=S1405-55462018000300881
Tags: Add Tag
No Tags, Be the first to tag this record!
id oai:scielo:S1405-55462018000300881
record_format ojs
spelling oai:scielo:S1405-554620180003008812020-02-04Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic ApproachSierra Martínez,Luz MarinaCobos,Carlos AlbertoCorrales Muñoz,Juan CarlosRojas Curieux,TulioHerrera-Viedma,EnriquePeluffo-Ordóñez,Diego Hernán Part of speech tagger Nasa Yuwe language tagged corpus harmony search global-best harmony search hill climbing tabu memory Abstract: Nasa Yuwe is the language of the Nasa indigenous community in Colombia. It is currently threatened with extinction. In this regard, a range of computer science solutions have been developed to the teaching and revitalization of the language. One of the most suitable approaches is the construction of a Part-Of-Speech Tagging (POST), which encourages the analysis and advanced processing of the language. Nevertheless, for Nasa Yuwe no tagged corpus exists, neither is there a POS Tagger and no related works have been reported. This paper therefore concentrates on building a linguistic corpus tagged for the Nasa Yuwe language and generating the first tagging application for Nasa Yuwe. The main results and findings are 1) the process of building the Nasa Yuwe corpus, 2) the tagsets and tagged sentences, as well as the statistics associated with the corpus, 3) results of two experiments to evaluate several POS Taggers (a Random tagger, three versions of HSTAGger, a tagger based on the harmony search metaheuristic, and three versions of a memetic algorithm GBHS Tagger, based on Global-Best Harmony Search (GBHS), Hill Climbing and an explicit Tabu memory, which obtained the best results in contrast with the other methods considered over the Nasa Yuwe language corpus.info:eu-repo/semantics/openAccessInstituto Politécnico Nacional, Centro de Investigación en ComputaciónComputación y Sistemas v.22 n.3 20182018-09-01info:eu-repo/semantics/articletext/htmlhttp://www.scielo.org.mx/scielo.php?script=sci_arttext&pid=S1405-55462018000300881en10.13053/cys-22-3-3018
institution SCIELO
collection OJS
country México
countrycode MX
component Revista
access En linea
databasecode rev-scielo-mx
tag revista
region America del Norte
libraryname SciELO
language English
format Digital
author Sierra Martínez,Luz Marina
Cobos,Carlos Alberto
Corrales Muñoz,Juan Carlos
Rojas Curieux,Tulio
Herrera-Viedma,Enrique
Peluffo-Ordóñez,Diego Hernán
spellingShingle Sierra Martínez,Luz Marina
Cobos,Carlos Alberto
Corrales Muñoz,Juan Carlos
Rojas Curieux,Tulio
Herrera-Viedma,Enrique
Peluffo-Ordóñez,Diego Hernán
Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
author_facet Sierra Martínez,Luz Marina
Cobos,Carlos Alberto
Corrales Muñoz,Juan Carlos
Rojas Curieux,Tulio
Herrera-Viedma,Enrique
Peluffo-Ordóñez,Diego Hernán
author_sort Sierra Martínez,Luz Marina
title Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
title_short Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
title_full Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
title_fullStr Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
title_full_unstemmed Building a Nasa Yuwe Language Corpus and Tagging with a Metaheuristic Approach
title_sort building a nasa yuwe language corpus and tagging with a metaheuristic approach
description Abstract: Nasa Yuwe is the language of the Nasa indigenous community in Colombia. It is currently threatened with extinction. In this regard, a range of computer science solutions have been developed to the teaching and revitalization of the language. One of the most suitable approaches is the construction of a Part-Of-Speech Tagging (POST), which encourages the analysis and advanced processing of the language. Nevertheless, for Nasa Yuwe no tagged corpus exists, neither is there a POS Tagger and no related works have been reported. This paper therefore concentrates on building a linguistic corpus tagged for the Nasa Yuwe language and generating the first tagging application for Nasa Yuwe. The main results and findings are 1) the process of building the Nasa Yuwe corpus, 2) the tagsets and tagged sentences, as well as the statistics associated with the corpus, 3) results of two experiments to evaluate several POS Taggers (a Random tagger, three versions of HSTAGger, a tagger based on the harmony search metaheuristic, and three versions of a memetic algorithm GBHS Tagger, based on Global-Best Harmony Search (GBHS), Hill Climbing and an explicit Tabu memory, which obtained the best results in contrast with the other methods considered over the Nasa Yuwe language corpus.
publisher Instituto Politécnico Nacional, Centro de Investigación en Computación
publishDate 2018
url http://www.scielo.org.mx/scielo.php?script=sci_arttext&pid=S1405-55462018000300881
work_keys_str_mv AT sierramartinezluzmarina buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
AT coboscarlosalberto buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
AT corralesmunozjuancarlos buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
AT rojascurieuxtulio buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
AT herreraviedmaenrique buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
AT peluffoordonezdiegohernan buildinganasayuwelanguagecorpusandtaggingwithametaheuristicapproach
_version_ 1756225795061186560