Leveraging Group Relative Policy Optimization to Advance Large Language Models In Traditional Chinese Medicine
Abstract:Traditional Chinese Medicine (TCM) presents a rich and structurally distinctive knowledge system that challenges typical functions of giant language fashions (LLMs). Although previous TCM-particular LLMs have shown progress by means of supervised nice-tuning, they typically face limitations in alignment, knowledge high quality, and analysis consistency. On this examine, we introduce Ladder-base, the first TCM-focused LLM educated with Group Relative Policy Optimization (GRPO), a reinforcement learning technique that improves reasoning and factual consistency by optimizing response selection primarily based on intra-group comparisons. Ladder-base is constructed upon the Qwen2.5-7B-Instruct foundation model and skilled exclusively on the textual subset of the TCM-Ladder benchmark, using 80 % of the info for training and the remaining 20 percent cut up evenly between validation and test sets. Through standardized evaluation, Ladder-base demonstrates superior performance throughout multiple reasoning metrics when compared to both state-of-the-art basic-goal LLMs equivalent to GPT-4, Gemini 2.5, Claude 3, and Qwen3 and area-particular TCM models together with BenTsao, HuatuoGPT2, and Zhongjing. These findings suggest that GRPO gives an effective and environment friendly technique for aligning LLMs with professional-level reasoning in traditional medical domains and helps the development of reliable and clinically grounded TCM artificial intelligence systems.