UPM Institutional Repository

Evaluating machine translation of the Shan Hai Jing: an MQM-based analysis of google translate vs. ChatGPT with prompting effects


Citation

Duan, Wenqi and Ng, Chwee Fang and Abdul Halim, Hazlina and Zhang, Zhongming (2026) Evaluating machine translation of the Shan Hai Jing: an MQM-based analysis of google translate vs. ChatGPT with prompting effects. GEMA Online Journal of Language Studies, 26 (2). pp. 404-426. ISSN 1675-8021; eISSN: 2550-2131

Abstract

Culturally dense classical texts pose persistent challenges for machine translation, particularly in reconstructing compressed semantic hierarchies and culture-specific references. Although Neural Machine Translation (NMT) and Large Language Models (LLMs) have substantially improved fluency and contextual coherence, previous studies have given limited attention to the evaluation of their performance on culturally embedded classical texts using MQM-based human evaluation alongside automatic metrics. Focusing on the English translation of the Shan Hai Jing, a culturally dense and semantically complex classical Chinese text, this study investigates whether different translation systems produce distinct error patterns in culturally compressed contexts and whether prompting strategies influence translation performance. Selected textual segments were translated using Google Translate and ChatGPT under minimal and enriched prompting strategies. Translation quality was assessed through MQM-based human evaluation alongside several automatic metrics (BLEU, chrF, BERTScore, and COMET-Kiwi). MQM analysis reveals clear differences in error patterns across systems: NMT outputs show a higher incidence of high-severity mistranslations, whereas LLM outputs tend to exhibit semantic generalisation and shifts in cultural references. By contrast, automatic metrics show limited differentiation in system rankings, with no significant main effect of system observed. Prompt enrichment does not produce consistent quality improvements and occasionally increases semantic drift. These findings suggest that translation quality in culturally compressed texts may be better interpreted through structural error patterns across MQM dimensions rather than metric-based rankings alone. Evaluation sensitivity appears to be shaped by text type, and increased prompt complexity does not necessarily enhance semantic precision in classical translation tasks.


Download File

[img] Text
126330.pdf - Published Version

Download (1MB)
Official URL or Download Paper: https://ejournal.ukm.my/gema/article/view/99954

Additional Metadata

Item Type: Article
Subject: Language and Linguistics
Subject: Linguistics and Language
Subject: Literature and Literary Theory
Divisions: Faculty of Modern Language and Communication
DOI Number: https://doi.org/10.17576/gema-2026-2602-08
Publisher: Penerbit Universiti Kebangsaan Malaysia
Keywords: Neural Machine Translation (NMT); Large Language Models (LLMs); Multidimensional Quality Metrics (MQM); Shan Hai Jing; Prompt strategies
Sustainable Development Goals (SDGs): SDG 9: Industry, Innovation and Infrastructure
Depositing User: Ms. Siti Radziah Mohamed@mahmod
Date Deposited: 16 Jul 2026 08:24
Last Modified: 16 Jul 2026 08:24
Altmetrics: http://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.17576/gema-2026-2602-08
URI: http://psasir.upm.edu.my/id/eprint/126330
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item