Megadose Built for builders and researchers.

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

· ArXiv · AI/CL/LG ·
PIHF leaves the model weights fixed and shifts learning into a reviewed, versioned language policy.

The paper describes a loop where a language-model critic and a clinical expert examine full reasoning and tool-use traces, identify recurring failures, and propose policy revisions. The expert keeps control over admitting changes and rolling them back. On ultra-rare-disease benchmarks, the resulting policy raised Recall@1 across one proprietary executor and three open-weight executors from 3B to 49B active parameters. Reported gains include 32.7 points for GPT-5.4 and 31.1 points for Qwen3.6-35B. ArXiv · AI/CL/LG's note

score 4

Categories: Research