Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
Wnuan’s staged adaptation lifted enterprise QA accuracy sharply, but it also cost the model some general ability.
The paper reports a three-stage post-training pipeline for proprietary enterprise knowledge: document-derived supervision, SFT with general-data replay, then RL on remaining errors. On WnuanBench, the 32B route rose from 52.76% acceptable answers before adaptation to 91.51% after RL. Residual-error sampling beat two comparison sampling setups under the same update budget. The authors also report a 5.17-point drop on general benchmarks, mainly in instruction following. HF Daily Papers' note
The paper reports a three-stage post-training pipeline for proprietary enterprise knowledge: document-derived supervision, SFT with general-data replay, then RL on remaining errors. On WnuanBench, the 32B route rose from 52.76% acceptable answers before adaptation to 91.51% after RL. Residual-error sampling beat two comparison sampling setups under the same update budget. The authors also report a 5.17-point drop on general benchmarks, mainly in instruction following. HF Daily Papers' note
score 4