ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
HF Daily Papers
· Sep 16, 2026
ProgramDistill benchmarks coding agents on inferred web-app behavior using replay-verified tasks mined from reference applications.
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
ArXiv · AI/CL/LG
· Sep 16, 2026
The paper argues model growth and recursive depth can alter scaling exponents, claiming GPT-3 13B performance at far lower compute.
Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
ArXiv · AI/CL/LG
· Sep 16, 2026
The work finds simple representation vectors can detect and analyze reward hacking across major open language models.
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
ArXiv · AI/CL/LG
· Sep 16, 2026
Andromeda 2 autonomously designs and tests drug formulations, outperforming optimization and wet-lab baselines on paclitaxel.
Higher-order pruning of experts in mixture-of-experts language models
ArXiv · AI/CL/LG
· Sep 16, 2026
HOPE improves mixture-of-experts pruning by modeling second-order interactions between experts across large language models.
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
ArXiv · AI/CL/LG
· Sep 16, 2026
ReFigBench evaluates multimodal coding agents on reconstructing scientific figures as editable PowerPoint artifacts.
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
ArXiv · AI/CL/LG
· Sep 16, 2026
The paper proposes generating and adapting model weights from live interaction data instead of relying only on prompts.
Using OCR Heads to Verbalize Image Semantics
ArXiv · AI/CL/LG
· Sep 16, 2026
The study finds OCR attention heads in VLMs act as general semantic verbalization heads across image tokens.
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
HF Daily Papers
· Sep 15, 2026
Emergence World stress-tests frontier-model agents over 16-day multi-agent simulations, exposing long-horizon failures across memory, tools, and state.
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
HF Daily Papers
· Sep 15, 2026
Zing-0.5 is a 5B world model for real-time playable generation controlled by keyboard actions and text.
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
ArXiv · AI/CL/LG
· Sep 15, 2026
ScienceBuddy combines an interactive research workspace with recursive harness and model improvement for scientific agents.
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
ArXiv · AI/CL/LG
· Sep 15, 2026
LimiX-2 introduces a contextual mechanism network for structured-data intelligence trained with context-conditional masked modeling.
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
ArXiv · AI/CL/LG
· Sep 15, 2026
A SWE-bench audit argues top coding-agent scores are statistically tied and proposes measuring model-scaffold interactions instead.
OPEN-1B: A Fully Auditable Training Run
ArXiv · AI/CL/LG
· Sep 15, 2026
OPEN-1B proposes a fully auditable model training run designed to make every operation and data sample independently verifiable.
Same Flow, Different Paths: Variance Reduction in Flow Matching
ArXiv · AI/CL/LG
· Sep 15, 2026
A flow matching analysis shows path choice can change SGD convergence even when the marginal objective is unchanged.
Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)
Techmeme
· Sep 15, 2026
SemiAnalysis benchmarks report Rubin NVL72 delivers up to 7x better token throughput per megawatt than Blackwell on a 1.6T DeepSeek model.
Atria Dawn: The Dawn of Agentic Superintelligence
HF Daily Papers
· Sep 14, 2026
Atria Dawn Preview is presented as a foundation agentic model for research and engineering workflows with strong benchmark results.
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
HF Daily Papers
· Sep 14, 2026
PhysBrain 1.5 unifies physical scene understanding, action generation, and future-state prediction in one autoregressive model.
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
ArXiv · AI/CL/LG
· Sep 14, 2026
K-Bench evaluates 125 LLM configurations on clinician-calibrated high-risk mental health conversations with strong clinician agreement.
Realtime-Venus: A full-duplex interaction system with asynchronous delegation
HF Daily Papers
· Sep 12, 2026
Realtime-Venus introduces two 9B full-duplex audio and audio-visual models for continuous real-time interaction.
StepAudio 3 Gen Technical Report
HF Daily Papers
· Sep 11, 2026
StepAudio 3 Gen uses discrete autoregressive RVQ token modeling for unified speech, music, sound effects, and mixed audio generation.
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
HF Daily Papers
· Sep 11, 2026
ZGCM-1 is a fully open 7B foundation model trained from scratch for math and agentic search with 256K context.
On the Navier–Stokes Millennium Prize Problem
OpenAI Blog
· Sep 8, 2026
OpenAI published an AI-generated Navier-Stokes solution with a Lean formalization, claiming a major mathematical breakthrough from its agentic research system.
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
HF Daily Papers
· Sep 10, 2026
Vidu S2 adds real-time 720p avatar generation and real-time stream editing with an online demo.
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
Google DeepMind
· Sep 8, 2026
DeepMind released AlphaGenome Atlas, a genome-wide predictive map for 9 billion human DNA variants.
Quoting Calif Research
Simon Willison
· Sep 10, 2026
Calif Research demonstrated WeWorm, a zero-click WeChat call worm built with AI-assisted exploit development.
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
HF Daily Papers
· Sep 10, 2026
T1 is a 122B MoE terminal agent trained with reinforcement learning for long-horizon shell tasks with verifier-based rewards.
Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
HF Daily Papers
· Sep 10, 2026
Pelican-Sim 1.0 is a world-model simulator for embodied AI that predicts future observations from visual context and robot actions.
Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
ArXiv · AI/CL/LG
· Sep 10, 2026
py-kvcache adds scheduler-aware NVMe KV-cache offload for vLLM, improving long-context serving performance under realistic workloads.
What OpenAI’s latest controversy tells us about the future of math
MIT Technology Review AI
· Sep 9, 2026
OpenAI says its agents solved a Millennium Prize Problem, a major mathematical capability claim now facing scrutiny.
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
HF Daily Papers
· Sep 9, 2026
NVIDIA details an open Nemotron-based pipeline for high-performing natural-language Olympiad proof generation without formal provers.
OpenAI says, while unlikely, it "cannot rule out that de-identified data derived" from Buckmaster's and Alpöge's use of its products helped improve its models (OpenAI)
Techmeme
· Sep 8, 2026
OpenAI published a Navier-Stokes research paper with a Lean formalized proof and disclosed possible training-data influence from product usage.
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
HF Daily Papers
· Sep 9, 2026
NCP-ArchPreview explores latent-space language modeling by jointly training next-token and next-concept prediction objectives.
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
HF Daily Papers
· Sep 9, 2026
LLaDA-UI applies block-wise diffusion language modeling to a 16.7B vision-language GUI agent for lower-latency action generation.
ExecCritic: Learn to Test, Test to Improve for Coding Agents
ArXiv · AI/CL/LG
· Sep 8, 2026
ExecCritic separates test generation from repair and trains coding agents with role-specific RL inside a fail-closed scaffold.
Training-Free Task Vectors for LLM Behavioral Control
ArXiv · AI/CL/LG
· Sep 8, 2026
Training-Free Task Vectors map activation steering into rank-one weight edits without fine-tuning.
Anthropic says Claude worked "largely autonomously" over 11 days to formalize the proof of Fermat's Last Theorem in the Lean programming language (Anthropic)
Techmeme
· Sep 4, 2026
Anthropic reports Claude autonomously completed a Lean formalization of Fermat's Last Theorem, demonstrating long-horizon theorem-proving capability.
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
HF Daily Papers
· Aug 31, 2026
Qwen3.8-Flash-Next details a sparse 125B MoE architecture that cuts training compute while matching larger predecessors.
Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
HN · Frontpage AI
· Aug 26, 2026
Qwen introduced a Flash-Next architecture focused on lower-cost inference, attracting major developer discussion around model efficiency.
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
HF Daily Papers
· Aug 25, 2026
LAION-BVD releases a 10-million-hour open video dataset aimed at large-scale multimodal pretraining across video, audio, and image tasks.
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
HF Daily Papers
· Aug 24, 2026
The Station multi-agent environment produced novel mathematical results across several AlphaEvolve-style construction problems.
Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels (Terry Chen/NVIDIA Technical Blog)
Techmeme
· Aug 21, 2026
Nvidia’s AVO agent system completed every ARC-AGI-3 public task, signaling a notable benchmark result for long-horizon autonomous coding agents.