Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Mobius-v0 separates stored knowledge from reasoning steps, then reports similar scores with less training data or faster inference.
The paper describes a shared FFN “Memory” that holds knowledge vectors and multiple self-attention “Reasoners” that query it through hidden states. A 7B model trained from scratch matched a 7B Transformer baseline using 62.6% of the baseline’s training data. Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, is reported to reach similar downstream scores with nearly 4x end-to-end inference speedup. Source: HF Daily Papers' note.
The paper describes a shared FFN “Memory” that holds knowledge vectors and multiple self-attention “Reasoners” that query it through hidden states. A 7B model trained from scratch matched a 7B Transformer baseline using 62.6% of the baseline’s training data. Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, is reported to reach similar downstream scores with nearly 4x end-to-end inference speedup. Source: HF Daily Papers' note.
score 6