Megadose AI progress, ranked and analyzed.

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

· HF Daily Papers ·
Wnuan’s staged adaptation lifted enterprise QA accuracy sharply, but it also cost the model some general ability.

The paper reports a three-stage post-training pipeline for proprietary enterprise knowledge: document-derived supervision, SFT with general-data replay, then RL on remaining errors. On WnuanBench, the 32B route rose from 52.76% acceptable answers before adaptation to 91.51% after RL. Residual-error sampling beat two comparison sampling setups under the same update budget. The authors also report a 5.17-point drop on general benchmarks, mainly in instruction following. HF Daily Papers' note

score 4

Categories: Research