Gascó Mora, GuillemRocha Sánchez, Martha AliciaSanchis Trilles, GermánAndrés Ferrer, JesúsCasacuberta Nolla, Francisco2014-01-292014-01-292012-04-23978-1-937284-19-0https://riunet.upv.es/handle/10251/35214Nowadays, there are large amounts of data available to train statistical machine translation systems. However, it is not clear whether all the training data actually help or not. A system trained on a subset of such huge bilingual corpora might outperform the use of all the bilingual data. This paper studies such issues by analysing two training data selection techniques: one based on approximating the probability of an indomain corpus; and another based on infrequent n-gram occurrence. Experimental results not only report significant improvements over random sentence selection but also an improvement over a system trained with the whole available data. Surprisingly, the improvements are obtained with just a small fraction of the data that accounts for less than 0.5% of the sentences. Afterwards, we show that a much larger room for improvement exists, although this is done under non-realistic conditions.Reserva de todos los derechosBilingual corporaTraining data selection techniquesProbability of an indomain corpusInfrequent n-gram occurrenceESTADISTICA E INVESTIGACION OPERATIVALENGUAJES Y SISTEMAS INFORMATICOSDoes more data always yield better translations?Comunicación en congresoAbierto