An experiment in evaluating the quality of translations

To lay the foundations for a systematic procedure that could be applied to any scientific translation, this experiment evaluates the error variances attributable to various sources inherent in a design in which discrete, randomly ordered sentences from translations are rated for intelligibility and for fidelity to the original. The procedure is applied to three human and three mechanical translations into English of four passages from a Russian work on cybernetics, yielding mean scores for the translations. Human and mechanical translations are clearly different in over-all quality, although substantial overlap is noted when individual sentences are considered. The procedure also clearly differentiates within sets of human translations and within sets of mechanical translations. Results from the two scales are highly correlated, and these in turn are highly correlated with reading times. A procedure in which highly intelligent monolingual raters (i.e., without knowledge of the foreign language) compare a test translation with a carefully prepared translation is found to be more reliable than one in which bilingual raters compare the English translation with the Russian original.

An experiment in evaluating the quality of translations | Litlas