Jeeves. Reasoning improves Jev-like decision models
Jeeves adds test-time reasoning to a Jev-style decision model and reports higher accuracy on several held-out and public benchmarks.
The repo describes a 9B Qwen3.5-based classifier trained with SFT and CISPO, plus a block-4 diffusion drafter and released training/eval code. It supports Jev-compatible yes/no, choice, and score requests, with optional reasoning before the final probability decision. The reported gains include 0.889 on its held-out test set versus 0.857 for Jev, and 0.935 on public JevBench tiers versus 0.866. The authors note limits: knowledge benchmarks trail Jev, full thinking has a 17.1s p90 latency on one H100, and the reasoning chains are not reliably interpretable. HN · GitHub AI's note
The repo describes a 9B Qwen3.5-based classifier trained with SFT and CISPO, plus a block-4 diffusion drafter and released training/eval code. It supports Jev-compatible yes/no, choice, and score requests, with optional reasoning before the final probability decision. The reported gains include 0.889 on its held-out test set versus 0.857 for Jev, and 0.935 on public JevBench tiers versus 0.866. The authors note limits: knowledge benchmarks trail Jev, full thinking has a 17.1s p90 latency on one H100, and the reasoning chains are not reliably interpretable. HN · GitHub AI's note
score 5
Discussions
- hn · 242 points · 95 comments