Papers
arxiv:2610.05966

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Published on Oct 5
ยท Submitted by
Wang
on Oct 7
Authors:
,
,
,
,
,
,
,
,

Abstract

Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.

Community

Paper submitter

HuatuoGPT-3 advances the HuatuoGPT line from medical data adaptation to training-paradigm innovation. Instead of following the conventional SFT-then-RL pipeline, it explores RL-only domain adaptation from base models through OnePO, using teacher outputs as temporary guidance while avoiding long-term imitation constraints. In this paper, we release an open-source medical LLM series whose 27B variant reaches 71.4 on HealthBench Professional. This makes HuatuoGPT-3 both a strong medical expert model family and a demonstration that specialized medical capabilities can be cultivated through direct policy optimization, reducing reliance on multi-stage supervised fine-tuning pipelines.

The HuatuoGPT series presents a progressive research trajectory for medical large language models, moving from medical dialogue alignment to domain adaptation, multimodal understanding, complex reasoning, and reinforcement-learning-based specialization.

  • HuatuoGPT first showed that combining ChatGPT-distilled data with real doctor-patient data and RLAIF can better align open LLMs with medical consultation needs [1].
  • HuatuoGPT-II then simplified medical domain adaptation by converting heterogeneous medical corpora and instruction data into a unified input-output format for one-stage training [2].
  • HuatuoGPT-Vision extended the series to medical multimodal intelligence by constructing PubMedVision and injecting large-scale medical visual knowledge into MLLMs [3].
  • HuatuoGPT-o1 further targeted complex medical reasoning through medical verifiable problems, verifier-guided reasoning-trajectory search, and reinforcement learning [4].
  • Building on this line, OnePO, accepted by ICML 2026, proposes direct one-stage policy optimization for SFT-free domain adaptation [5], and HuatuoGPT-3 scales this idea into an RL-only medical LLM series, marking a shift from data-centric medical adaptation toward training-paradigm innovation [6].

References

[1] Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, and Haizhou Li. HuatuoGPT, Towards Taming Language Model to Be a Doctor. EMNLP Findings, 2023.
[2] Junying Chen, Xidong Wang, Ke Ji, Anningzhe Gao, Feng Jiang, Shunian Chen, Hongbo Zhang, Dingjie Song, Wenya Xie, Chuyi Kong, Jianquan Li, Xiang Wan, Haizhou Li, and Benyou Wang. HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs. COLM, 2024.
[3] Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao, Shunian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, Guangjun Yu, Xiang Wan, and Benyou Wang. HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale. EMNLP, 2024.
[4] Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, and Benyou Wang. HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs. arXiv, 2024.
[5] Junying Chen, Xinyuan Xie, Ziniu Li, and Benyou Wang. OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation. ICML, 2026.
[6] Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, and Benyou Wang. HuatuoGPT-3: RL-Only Domain Adaptation from Base Models. arXiv, 2026.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.05966
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.05966 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.05966 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.