ISSN 0201-7385. ISSN 2074-6636
En Ru
ISSN 0201-7385. ISSN 2074-6636
Multi-Agent Assessment of Machine Translation Quality Across Subject Domains: A Comparative Analysis of Google Translate, Yandex Translate and DeepL

Multi-Agent Assessment of Machine Translation Quality Across Subject Domains: A Comparative Analysis of Google Translate, Yandex Translate and DeepL

Abstract

Object: this study presents a comparative analysis of the quality of machine translation (MT) produced by three leading systems (Google Translate, Yandex Translate, DeepL) across three specialized domains (medicine, law, and news articles). The analysis employs an original multi-agent approach based on the YandexGPT large language model.

Methods: the research material comprises 60 English-language texts (20 per domain), each translated by the three MT engines, resulting in a total of 180 translations for evaluation. A system of six virtual expert agents powered by YandexGPT was developed for assessment: a linguist, a semanticist, a terminologist, a pragmatist, a cultural adaptor, and a domain-specific expert (medical/legal/media discourse). Each agent evaluated translations on a 10-point scale. An agent-moderator then synthesized these individual assessments to produce a final score. Statistical analysis included ANOVA, Student's t-test, and Pearson/Spearman correlation analysis.

Findings: DeepL achieved the highest overall score (8.57±0.69), demonstrating particular strength in the legal domain (8.47). Google Translate unexpectedly excelled in the medical domain (8.80) and also led in media texts (8.59). Yandex Translate ranked third across all domains, exhibiting the highest score variability (7.91±1.06) and critical errors in legal translations (minimum score of 5.0). Statistically significant differences were found between the MT engines (p < 0.05). Correlation analysis of agent scores revealed high consistency between the linguist and semanticist (r = 0.82).

Conclusions: the proposed multi-agent approach provides explainable, multi-aspect evaluations. DeepL emerges as a versatile leader, particularly for legal translation; Google is preferable for medical and media texts; Yandex requires further development for specialized domains. These findings can inform post-editing strategies and guide the selection of MT systems for professional tasks.


PDF, ru

Received: 02/20/2026

Accepted: 05/05/2026

Accepted date: 26.05.2026

Keywords: machine translation, multi-agent systems, YandexGPT, quality assessment, Google Translate, DeepL, Yandex Translate, medical translation, legal translation

DOI: 10.55959/MSU2074-6636-22-2026-19-2-131-156

Available in the on-line version with: 22.07.2026

  • To cite this article:
Issue 2, 2026