Megadose AI progress, ranked and analyzed.

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

· ArXiv · AI/CL/LG ·
Deep Noir automates where and how hard to apply activation steering, then reports stronger classifier gains than baseline steering methods.

The paper says its framework combines Logit Lens convergence with causal head-level attribution to find steering parameters without manual selection. It reports spam-classification gains of 16.7 percentage points at 1B scale and 21 to 42 points across four 7-9B architectures. On SST-2 sentiment, it reports a 13.1-point gain with no code changes, while RepE without head masking did not beat baseline. The authors also warn that stronger steering creates a more predictable prompt-injection attack surface. ArXiv · AI/CL/LG's note

score 4

Categories: Research