Viktor O. Savostyanov
-
Multi-Agent Assessment of Machine Translation Quality Across Subject Domains: A Comparative Analysis of Google Translate, Yandex Translate and DeepLMoscow University Translation Studies Bulletin. 2026. Vol. 19. № 2. p.131-156read more57
-
Object: this study presents a comparative analysis of the quality of machine translation (MT) produced by three leading systems (Google Translate, Yandex Translate, DeepL) across three specialized domains (medicine, law, and news articles). The analysis employs an original multi-agent approach based on the YandexGPT large language model.
Methods: the research material comprises 60 English-language texts (20 per domain), each translated by the three MT engines, resulting in a total of 180 translations for evaluation. A system of six virtual expert agents powered by YandexGPT was developed for assessment: a linguist, a semanticist, a terminologist, a pragmatist, a cultural adaptor, and a domain-specific expert (medical/legal/media discourse). Each agent evaluated translations on a 10-point scale. An agent-moderator then synthesized these individual assessments to produce a final score. Statistical analysis included ANOVA, Student's t-test, and Pearson/Spearman correlation analysis.
Findings: DeepL achieved the highest overall score (8.57±0.69), demonstrating particular strength in the legal domain (8.47). Google Translate unexpectedly excelled in the medical domain (8.80) and also led in media texts (8.59). Yandex Translate ranked third across all domains, exhibiting the highest score variability (7.91±1.06) and critical errors in legal translations (minimum score of 5.0). Statistically significant differences were found between the MT engines (p < 0.05). Correlation analysis of agent scores revealed high consistency between the linguist and semanticist (r = 0.82).
Conclusions: the proposed multi-agent approach provides explainable, multi-aspect evaluations. DeepL emerges as a versatile leader, particularly for legal translation; Google is preferable for medical and media texts; Yandex requires further development for specialized domains. These findings can inform post-editing strategies and guide the selection of MT systems for professional tasks.
Keywords: machine translation, multi-agent systems, YandexGPT, quality assessment, Google Translate, DeepL, Yandex Translate, medical translation, legal translation
-

