pisco_log
banner

Multi-Dimensional Translation Quality Assessment of Astronomical Classics via Reasoning Large Language Models

Xiuwen Wang, Chiyu Pan

Abstract


Taking the astronomical classic 《授时历》 as a case study, this research tests the performance of LLMs as translation quality assessment tools. Based on a framework integrating House's functional-pragmatic principles and the MQM scale, two reasoning models,
Deepseek-R1 and Gemini-3.0-Pro, are introduced to verify the automated assessment capabilities of LLMs by calculating their consistency
with traditional automatic metrics and human assessments. The results indicate that: (1) The ranking capability of LLMs for translation quality surpasses that of traditional automatic metrics; (2) LLMs show a moderate correlation with human scoring in the semantic dimension but
perform poorly in pragmatic and textual dimensions. This study clarifies the performance of LLMs in the quality assessment of astronomical
classics, providing a reference for technological innovation in translation quality assessment.

Keywords


Large Language Models; Translation Quality Assessment; LLM-based Automated Assessment

Full Text:

PDF

Included Database


References


[1] Guo, D., Zhu, C., Yang, C., Wang, Z., Yu, X., Ganshin, I., ... & Zhou, S. (2025). Deepseek-r1: Incentivizing reasoning capability in llms

via reinforcement learning. arXiv. https://doi.org/10.48550/arXiv.2501.12948

[2] House, J. (1997). Translation quality assessment: A model revised. Narr.

[3] Kocmi, T., & Federmann, C. (2023). GEMBA-MQM: Detecting translation quality error spans with GPT-4. In Proceedings of the Eighth

Conference on Machine Translation (pp. 768–775). Association for Computational Linguistics.

[4] Lavie, A., & Agarwal, A. (2007). METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation (pp. 228–231). Association for Computational Linguistics.

[5] Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (pp. 311–318). Association for Computational

Linguistics.

[6] Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020

Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2685–2702). Association for Computational Linguistics.

[7] Reiss, K. (1971). Text types, translation types and translation assessment. In A. Chesterman (Ed.), Readings in translation theory (pp.

105–115). Oy Finn Lectura Ab.

[8] Sivin, N. (2009). Granting the seasons: The Chinese astronomical reform of 1280, with a study of its many dimensions and a translation

of its records. Springer Science+Business Media, LLC.

[9] Williams, M. (2004). Translation quality assessment: An argumentation-centred approach. University of Ottawa Press.

[10] Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). BERTScore: Evaluating text generation with BERT. arXiv. https://

arxiv.org/abs/1904.09675




DOI: http://dx.doi.org/10.70711/rcha.v4i6.9986

Refbacks

  • There are currently no refbacks.