Evaluating Google’s new TranslateGemma for English–Arabic machine translation: A corpus-based assessment using United Nations documents
DOI:
https://doi.org/10.53696/27753719.61411Keywords:
Machine translation, English translation, Arabic translation, Neural machine translation, Evaluation metrics, United Nations corpusAbstract
This study presents a systematic evaluation of Google’s newly released model TranslateGemma-4B for English-Arabic machine translation using official United Nations parallel documents. Despite the recent proliferation of large language models claiming multilingual competence, empirical assessments of translation quality for morphologically rich languages which are languages that exhibit many grammatical cases and inflected word forms such as Arabic remain limited, particularly when evaluated against professionally approved institutional reference translations. We assessed TranslateGemma in a zero-shot configuration (i.e., the prompt used contained only the task objective and no examples or demonstrations) on 10,000 sentence pairs drawn from the UN English–Arabic Parallel Corpus. The model was executed on Google Colab due to computational requirements exceeding local hardware capacity. Translation outputs were evaluated using five complementary automatic metrics: BLEU (6.95), chrF++ (33.28), METEOR (20.90), BERTScore (74.21), and COMET (71.60). Additionally, we implemented diagnostic heuristics to detect omissions, hallucinations (i.e., inventing irrelevant texts), digit mismatches, and terminology inconsistencies. Results indicate that while TranslateGemma achieves moderate semantic similarity scores on neural metrics, it exhibits substantial deficiencies in surface-form accuracy and institutional terminology consistency. Sentence length ratio analysis revealed systematic under-translation patterns, with 16.86% of outputs flagged for potential omission. Digit mismatch errors affected 24.48% of the corpus, raising concerns for high-stakes translation contexts. Terminology consistency analysis using an expanded UN glossary indicated that 7.95% of sentences containing frozen institutional terms failed to preserve standard Arabic equivalents. These findings demonstrate that TranslateGemma, while showing promise in capturing broad semantic adequacy, requires significant refinement before deployment in institutional translation workflows where precision and terminological fidelity are paramount.
Downloads
References
Abdelaal, N., & Al Sawi, I. (2025). A comparative evaluation of machine translation vs. human translation for legal texts: A case study of translations between English to Arabic. Comparative Legilinguistics, 63, 186–223. https://doi.org/10.14746/cl.2025.63.1
Abu-Ayyash, E. A. (2017). Errors and non-errors in English-Arabic machine translation of gender-bound constructs in technical texts. Procedia Computer Science, 117, 73–80. https://doi.org/10.1016/j.procs.2017.10.095 DOI: https://doi.org/10.1016/j.procs.2017.10.095
Ahmed, O., Tawfik, K., & Tohamy, M. (2025). Figurative language as a semantic barrier in the Arabic/English translation of United Nations General Assembly speeches. British Journal of Translation Linguistics and Literature, 5(2), 55–85. https://doi.org/10.54848/mvgwe536 DOI: https://doi.org/10.54848/mvgwe536
Alghamdi, E. A., Zakraoui, J., & Abanmy, F. A. (2024). Domain adaptation for Arabic machine translation: Financial texts as a case study. Applied Sciences, 14(16), 7088. https://doi.org/10.3390/app14167088 DOI: https://doi.org/10.3390/app14167088
Al-Khalifa, H., Al-Khalefah, K., & Haroon, H. (2024). Error analysis of pretrained language models (PLMs) in English-to-Arabic machine translation. Human-Centric Intelligent Systems, 4(2), 206–219. https://doi.org/10.1007/s44230-024-00061-7 DOI: https://doi.org/10.1007/s44230-024-00061-7
Alkhatib, M., & Shaalan, K. (2018). The key challenges for Arabic machine translation. In K. Shaalan, A. E. Hassanien, & F. Tolba (Eds.), Studies in computational intelligence (pp. 139–156). Springer. https://doi.org/10.1007/978-3-319-67056-0_8 DOI: https://doi.org/10.1007/978-3-319-67056-0_8
Alqudsi, A., Omar, N., & Shaker, K. (2012). Arabic machine translation: A survey. Artificial Intelligence Review, 42(4), 549–572. https://doi.org/10.1007/s10462-012-9351-1 DOI: https://doi.org/10.1007/s10462-012-9351-1
Altakhaineh, A. R. M., Alghathian, G. A., & Jarrah, M. M. (2024). A comparative study of accuracy in human vs. AI translation of legal documents into Arabic. International Journal of Language and Law (JLL). https://doi.org/10.14762/jll.2025.063
Alwazna, R. Y. (2024). The use of automation in the rendition of certain articles of the Saudi Commercial Law into English: A post-editing-based comparison of five machine translation systems. Frontiers in Artificial Intelligence, 6. https://doi.org/10.3389/frai.2023.1282020 DOI: https://doi.org/10.3389/frai.2023.1282020
Ameur, M. S. H., Meziane, F., & Guessoum, A. (2020). Arabic machine translation: A survey of the latest trends and challenges. Computer Science Review, 38, 100305. https://doi.org/10.1016/j.cosrev.2020.100305 DOI: https://doi.org/10.1016/j.cosrev.2020.100305
Badah, A. M., Khalaf, C. N., Dwaikat, F. J., & AlQbailat, N. (2025). The usage of artificial intelligence in legal translation: Bridging the gap between law and language. Ampersand, 16, 100248. https://doi.org/10.1016/j.amper.2025.100248 DOI: https://doi.org/10.1016/j.amper.2025.100248
Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. Association for Computational Linguistics. https://aclanthology.org/W05-0909/
Bechara, H., Manohara, K., & Jankin, S. (2024). Creating and evaluating a multilingual corpus of UN General Assembly debates. In Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1). European Association for Machine Translation. https://aclanthology.org/2024.eamt-1.52/
Bouamor, H., Alshikhabobakr, H., Mohit, B., & Oflazer, K. (2014). A human judgement corpus and a metric for Arabic MT evaluation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 207–213). Association for Computational Linguistics. https://doi.org/10.3115/v1/d14-1026 DOI: https://doi.org/10.3115/v1/D14-1026
Diab, N. (2021). Out of the BLEU: An error analysis of statistical and neural machine translation of WikiHow articles from English into Arabic. CDELT Occasional Papers in the Development of English Education, 75(1), 181–211. https://doi.org/10.21608/opde.2021.208437 DOI: https://doi.org/10.21608/opde.2021.208437
Guzmán, F., Bouamor, H., Baly, R., & Habash, N. (2016). Machine translation evaluation for Arabic using morphologically-enriched embeddings. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers. Association for Computational Linguistics. https://aclanthology.org/C16-1132/
Khasawneh, R. R., & Alsharif, B. B. (2025). A comparative study of AI-powered tools for Arabic-English and English-Arabic translation. Journal of Language Teaching and Research, 16(6), 2025–2035. https://doi.org/10.17507/jltr.1606.24 DOI: https://doi.org/10.17507/jltr.1606.24
Mohammed, T. A. S. (2025). Evaluating translation quality: A qualitative and quantitative assessment of machine and LLM-driven Arabic–English translations. Information, 16(6), 440. https://doi.org/10.3390/info16060440 DOI: https://doi.org/10.3390/info16060440
Moneus, A. M., & Sahari, Y. (2024). Artificial intelligence and human translation: A contrastive study based on legal texts. Heliyon, 10(6), e28106. https://doi.org/10.1016/j.heliyon.2024.e28106 DOI: https://doi.org/10.1016/j.heliyon.2024.e28106
Nassar, H. (2025). Challenges of post-editing in English to Arabic machine translation of technical texts: A study of technological and linguistic barriers. International Journal of Linguistics Literature & Translation, 8(4), 01–15. https://doi.org/10.32996/ijllt.2025.8.4.1 DOI: https://doi.org/10.32996/ijllt.2025.8.4.1
Popović, M. (2015). chrF: Character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation (pp. 392–395). Association for Computational Linguistics. https://doi.org/10.18653/v1/W15-3049 DOI: https://doi.org/10.18653/v1/W15-3049
Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2685–2702). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.213 DOI: https://doi.org/10.18653/v1/2020.emnlp-main.213
Hadla, L. S., Hailat, T. M., Al-Kabi, M. N. (2015). Comparative study between METEOR and BLEU methods of MT: Arabic into English translation as a case study. International Journal of Advanced Computer Science and Applications, 6(11). https://doi.org/10.14569/ijacsa.2015.061128 DOI: https://doi.org/10.14569/IJACSA.2015.061128
Sadiq, S. (2025). Evaluating English-Arabic translation: Human translators vs. Google Translate and ChatGPT. Journal of Languages and Translation, 12(1), 67–95. https://doi.org/10.21608/jltmin.2025.423147 DOI: https://doi.org/10.21608/jltmin.2025.423147
Shquier, M. M. A., & Sembok, T. M. T. (2008). Word agreement and ordering in English-Arabic machine translation. In 2008 International Symposium on Information Technology (pp. 1–10). IEEE. https://doi.org/10.1109/itsim.2008.4631625 DOI: https://doi.org/10.1109/ITSIM.2008.4631625
Vilar, D., & Black, K. (2026, January 16). TranslateGemma: A new suite of open translation models. Google. https://blog.google/innovation-and-ai/technology/developers-tools/translategemma/
Zakraoui, J., Saleh, M., Al-Maadeed, S., & Alja'am, J. M. (2020). Evaluation of Arabic to English machine translation systems. In 2020 11th International Conference on Information and Communication Systems (ICICS) (pp. 185–190). IEEE. https://doi.org/10.1109/icics49469.2020.239518 DOI: https://doi.org/10.1109/ICICS49469.2020.239518
Zakraoui, J., Saleh, M., Al-Maadeed, S., & Alja'am, J. M. (2021). Arabic machine translation: A survey with challenges and future directions. IEEE Access, 9, 161445–161468. https://doi.org/10.1109/access.2021.3132488 DOI: https://doi.org/10.1109/ACCESS.2021.3132488
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr
Ziemski, M., Junczys-Dowmunt, M., & Pouliquen, B. (2016). The United Nations Parallel Corpus v1.0. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). European Language Resources Association. https://aclanthology.org/L16-1561/
Zouidine, M., & Khalil, M. (2025). Large language models for Arabic sentiment analysis and machine translation. Engineering Technology & Applied Science Research, 15(2), 20737–20742. https://doi.org/10.48084/etasr.9584 DOI: https://doi.org/10.48084/etasr.9584
Downloads
Published
How to Cite
License
Copyright (c) 2026 Yassine El Rhaffouli, Hicham Boughaba

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
The authors certify that:
- if the manuscript is co-authored, they are authorized by their co-authors to enter into these arrangements.
- the work described has not been formally published before in a registered ISSN or ISBN media, except in the form of an abstract or as part of a published lecture, review, or thesis.
- it is not under consideration for publication elsewhere,
- its publication has been approved by all the author(s) and by the responsible authorities – tacitly or explicitly – of the institutes where the work has been carried out.
- they secure the right to reproduce any material that has already been published or copyrighted elsewhere (it does not infringe on the rights of others).
- they agree to license and copyright agreement.
All articles published are licensed under Creative Commons Attribution-NonCommercial 4.0 International License.
- Authors retain copyright and other proprietary rights related to the article.
- Authors retain the right and are permitted to use the substance of the article in their own future works, including lectures and books.
- Authors grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution-NonCommercial License (CC BY-NC 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post or self-archive their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.




