QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code
SFT after trading-code pretraining produced the strongest executable-strategy gains.
The paper tests language models on generating Backtrader strategies from natural-language requests. Continued pretraining improved single-turn judge pass rates, but by itself hurt final repaired agent performance. Adding supervised fine-tuning on agent-validated request-to-code pairs raised Qwen3.6-35B-A3B to 58.2% judge pass, 83.5% successful backtests, and 79.5% final agentic success after up to 10 turns. The authors also report that specialization damaged structured tool-calling behavior, and recovery SFT fixed formatting without restoring base repository-level agent performance. HF Daily Papers' note
The paper tests language models on generating Backtrader strategies from natural-language requests. Continued pretraining improved single-turn judge pass rates, but by itself hurt final repaired agent performance. Adding supervised fine-tuning on agent-validated request-to-code pairs raised Qwen3.6-35B-A3B to 58.2% judge pass, 83.5% successful backtests, and 79.5% final agentic success after up to 10 turns. The authors also report that specialization damaged structured tool-calling behavior, and recovery SFT fixed formatting without restoring base repository-level agent performance. HF Daily Papers' note
score 4