Large Language Models for Zhuang Medicine Translation: Translation Quality Assessment and A Human–AI Collaborative Translation Pathway
Tingting Long
Youjiang Medical University for Nationalities, Baise, Guangxi, China.
Meijuan Zhao *
Youjiang Medical University for Nationalities, Baise, Guangxi, China.
*Author to whom correspondence should be addressed.
Abstract
Zhuang Medicine, an important component of China’s ethnic medicine system, embodies distinctive medical theories and cultural knowledge whose accurate English translation is essential for international communication and global dissemination. However, existing studies have primarily focused on terminology translation, while systematic empirical evaluations of Large Language Models (LLMs) for Zhuang Medicine text translation remain limited. Using Practical Zhuang Medicine Techniques as the source corpus, this study constructed a test dataset of 240 representative sentences and systematically evaluated the Chinese–English translation performance of five representative LLMs (ChatGPT-5.2, Claude Sonnet 4.5, Gemini 2.5 Pro, DeepSeek-V3.2, and ERNIE 5.0) using the COMET metric, supplemented by statistical tests and qualitative analyses. The results indicate that DeepSeek-V3.2 achieved the highest mean COMET score (0.0249), followed by Claude Sonnet 4.5 (0.0218) and Gemini 2.5 Pro (0.0215). However, the Nemenyi post-hoc test indicated that statistically significant differences were only observed between ERNIE 5.0 and DeepSeek-V3.2, and between ERNIE 5.0 and Claude Sonnet 4.5. Qualitative analysis further demonstrates that translation quality is strongly dependent on text type: all models performed consistently well on texts with standardised terminology and low cultural specificity but encountered common difficulties in translating culturally embedded concepts, non-standardised terminology, and information-dense clinical descriptions. Based on these findings, this study proposes a three-tier Human–AI Collaborative Translation pathway comprising resource construction, model adaptation, and human quality assurance, offering practical insights into intelligent translation and human–AI collaborative workflows in low-resource specialised domains.
Keywords: Large language models (LLMs), machine translation (MT), translation quality assessment (TQA), COMET, Zhuang medicine texts, human-AI collaborative translation