Multi-Dimensional Translation Quality Assessment of Astronomical Classics via Reasoning Large Language Models
Abstract
Deepseek-R1 and Gemini-3.0-Pro, are introduced to verify the automated assessment capabilities of LLMs by calculating their consistency
with traditional automatic metrics and human assessments. The results indicate that: (1) The ranking capability of LLMs for translation quality surpasses that of traditional automatic metrics; (2) LLMs show a moderate correlation with human scoring in the semantic dimension but
perform poorly in pragmatic and textual dimensions. This study clarifies the performance of LLMs in the quality assessment of astronomical
classics, providing a reference for technological innovation in translation quality assessment.
Keywords
Full Text:
PDFReferences
[1] Guo, D., Zhu, C., Yang, C., Wang, Z., Yu, X., Ganshin, I., ... & Zhou, S. (2025). Deepseek-r1: Incentivizing reasoning capability in llms
via reinforcement learning. arXiv. https://doi.org/10.48550/arXiv.2501.12948
[2] House, J. (1997). Translation quality assessment: A model revised. Narr.
[3] Kocmi, T., & Federmann, C. (2023). GEMBA-MQM: Detecting translation quality error spans with GPT-4. In Proceedings of the Eighth
Conference on Machine Translation (pp. 768–775). Association for Computational Linguistics.
[4] Lavie, A., & Agarwal, A. (2007). METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation (pp. 228–231). Association for Computational Linguistics.
[5] Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (pp. 311–318). Association for Computational
Linguistics.
[6] Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020
Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2685–2702). Association for Computational Linguistics.
[7] Reiss, K. (1971). Text types, translation types and translation assessment. In A. Chesterman (Ed.), Readings in translation theory (pp.
105–115). Oy Finn Lectura Ab.
[8] Sivin, N. (2009). Granting the seasons: The Chinese astronomical reform of 1280, with a study of its many dimensions and a translation
of its records. Springer Science+Business Media, LLC.
[9] Williams, M. (2004). Translation quality assessment: An argumentation-centred approach. University of Ottawa Press.
[10] Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). BERTScore: Evaluating text generation with BERT. arXiv. https://
arxiv.org/abs/1904.09675
DOI: http://dx.doi.org/10.70711/rcha.v4i6.9986
Refbacks
- There are currently no refbacks.