Přístupnostní navigace
E-application
Search Search Close
Publication result detail
GRÉZL, F.; EGOROVA, E.; KARAFIÁT, M.
Original Title
Study of Large Data Resources for Multilingual Training and System Porting
English Title
Type
Paper in proceedings (conference paper)
Original Abstract
This study investigates the behavior of a feature extraction neural network model trained on a large amount of single language data("source language") on a set of under-resourced target languages. The coverage of the source language acoustic space was changedin two ways: (1) by changing the amount of training data and (2) by altering the level of detail of acoustic units (by changingthe triphone clustering). We observe the effect of these changes on the performance on target language in two scenarios: (1) thesource-language NNs were used directly, (2) NNs were first ported to target language.The results show that increasing coverage as well as level of detail on the source language improves the target language systemperformance in both scenarios. For the first one, both source language characteristic have about the same effect. For the secondscenario, the amount of data in source language is more important than the level of detail.The possibility to include large data into multilingual training set was also investigated. Our experiments point out possiblerisk of over-weighting the NNs towards the source language with large data. This degrades the performance on part of the targetlanguages, compared to the setting where the amounts of data per language are balanced.
English abstract
Keywords
Stacked Bottle-Neck; feature extraction; multilingual training; large data; Fisher database
Key words in English
Authors
RIV year
2018
Released
09.05.2016
Publisher
Elsevier Science
Location
Yogyakarta
Book
Procedia Computer Science
ISBN
1877-0509
Periodical
Volume
2016
Number
81
State
Kingdom of the Netherlands
Pages from
15
Pages to
22
Pages count
8
URL
http://www.sciencedirect.com/science/article/pii/S1877050916300382
BibTex
@inproceedings{BUT130953, author="František {Grézl} and Ekaterina {Egorova} and Martin {Karafiát}", title="Study of Large Data Resources for Multilingual Training and System Porting", booktitle="Procedia Computer Science", year="2016", journal="Procedia Computer Science", volume="2016", number="81", pages="15--22", publisher="Elsevier Science", address="Yogyakarta", doi="10.1016/j.procs.2016.04.024", issn="1877-0509", url="http://www.sciencedirect.com/science/article/pii/S1877050916300382" }
Documents
grezl_sltu2016_08-8028