Megadose AI progress, ranked and analyzed.

Research

Latest Research on Megadose. AI news ranked, decayed, deduped.

42 recent items

  1. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
    HF Daily Papers ·
    ProgramDistill benchmarks coding agents on inferred web-app behavior using replay-verified tasks mined from reference applications.
  2. How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
    ArXiv · AI/CL/LG ·
    The paper argues model growth and recursive depth can alter scaling exponents, claiming GPT-3 13B performance at far lower compute.
  3. Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
    ArXiv · AI/CL/LG ·
    The work finds simple representation vectors can detect and analyze reward hacking across major open language models.
  4. Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
    ArXiv · AI/CL/LG ·
    Andromeda 2 autonomously designs and tests drug formulations, outperforming optimization and wet-lab baselines on paclitaxel.
  5. Higher-order pruning of experts in mixture-of-experts language models
    ArXiv · AI/CL/LG ·
    HOPE improves mixture-of-experts pruning by modeling second-order interactions between experts across large language models.
  6. ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
    ArXiv · AI/CL/LG ·
    ReFigBench evaluates multimodal coding agents on reconstructing scientific figures as editable PowerPoint artifacts.
  7. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
    ArXiv · AI/CL/LG ·
    The paper proposes generating and adapting model weights from live interaction data instead of relying only on prompts.
  8. Using OCR Heads to Verbalize Image Semantics
    ArXiv · AI/CL/LG ·
    The study finds OCR attention heads in VLMs act as general semantic verbalization heads across image tokens.
  9. Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
    HF Daily Papers ·
    Emergence World stress-tests frontier-model agents over 16-day multi-agent simulations, exposing long-horizon failures across memory, tools, and state.
  10. Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
    HF Daily Papers ·
    Zing-0.5 is a 5B world model for real-time playable generation controlled by keyboard actions and text.
  11. ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
    ArXiv · AI/CL/LG ·
    ScienceBuddy combines an interactive research workspace with recursive harness and model improvement for scientific agents.
  12. LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
    ArXiv · AI/CL/LG ·
    LimiX-2 introduces a contextual mechanism network for structured-data intelligence trained with context-conditional masked modeling.
  13. Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
    ArXiv · AI/CL/LG ·
    A SWE-bench audit argues top coding-agent scores are statistically tied and proposes measuring model-scaffold interactions instead.
  14. OPEN-1B: A Fully Auditable Training Run
    ArXiv · AI/CL/LG ·
    OPEN-1B proposes a fully auditable model training run designed to make every operation and data sample independently verifiable.
  15. Same Flow, Different Paths: Variance Reduction in Flow Matching
    ArXiv · AI/CL/LG ·
    A flow matching analysis shows path choice can change SGD convergence even when the marginal objective is unchanged.
  16. Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)
    Techmeme ·
    SemiAnalysis benchmarks report Rubin NVL72 delivers up to 7x better token throughput per megawatt than Blackwell on a 1.6T DeepSeek model.
  17. Atria Dawn: The Dawn of Agentic Superintelligence
    HF Daily Papers ·
    Atria Dawn Preview is presented as a foundation agentic model for research and engineering workflows with strong benchmark results.
  18. PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
    HF Daily Papers ·
    PhysBrain 1.5 unifies physical scene understanding, action generation, and future-state prediction in one autoregressive model.
  19. K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
    ArXiv · AI/CL/LG ·
    K-Bench evaluates 125 LLM configurations on clinician-calibrated high-risk mental health conversations with strong clinician agreement.
  20. Realtime-Venus: A full-duplex interaction system with asynchronous delegation
    HF Daily Papers ·
    Realtime-Venus introduces two 9B full-duplex audio and audio-visual models for continuous real-time interaction.
  21. StepAudio 3 Gen Technical Report
    HF Daily Papers ·
    StepAudio 3 Gen uses discrete autoregressive RVQ token modeling for unified speech, music, sound effects, and mixed audio generation.
  22. ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
    HF Daily Papers ·
    ZGCM-1 is a fully open 7B foundation model trained from scratch for math and agentic search with 256K context.
  23. On the Navier–Stokes Millennium Prize Problem
    OpenAI Blog ·
    OpenAI published an AI-generated Navier-Stokes solution with a Lean formalization, claiming a major mathematical breakthrough from its agentic research system.
  24. Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
    HF Daily Papers ·
    Vidu S2 adds real-time 720p avatar generation and real-time stream editing with an online demo.
  25. AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
    Google DeepMind ·
    DeepMind released AlphaGenome Atlas, a genome-wide predictive map for 9 billion human DNA variants.
  26. Quoting Calif Research
    Simon Willison ·
    Calif Research demonstrated WeWorm, a zero-click WeChat call worm built with AI-assisted exploit development.
  27. T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
    HF Daily Papers ·
    T1 is a 122B MoE terminal agent trained with reinforcement learning for long-horizon shell tasks with verifier-based rewards.
  28. Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
    HF Daily Papers ·
    Pelican-Sim 1.0 is a world-model simulator for embodied AI that predicts future observations from visual context and robot actions.
  29. Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
    ArXiv · AI/CL/LG ·
    py-kvcache adds scheduler-aware NVMe KV-cache offload for vLLM, improving long-context serving performance under realistic workloads.
  30. What OpenAI’s latest controversy tells us about the future of math
    MIT Technology Review AI ·
    OpenAI says its agents solved a Millennium Prize Problem, a major mathematical capability claim now facing scrutiny.
  31. An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
    HF Daily Papers ·
    NVIDIA details an open Nemotron-based pipeline for high-performing natural-language Olympiad proof generation without formal provers.
  32. OpenAI says, while unlikely, it "cannot rule out that de-identified data derived" from Buckmaster's and Alpöge's use of its products helped improve its models (OpenAI)
    Techmeme ·
    OpenAI published a Navier-Stokes research paper with a Lean formalized proof and disclosed possible training-data influence from product usage.
  33. NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
    HF Daily Papers ·
    NCP-ArchPreview explores latent-space language modeling by jointly training next-token and next-concept prediction objectives.
  34. LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
    HF Daily Papers ·
    LLaDA-UI applies block-wise diffusion language modeling to a 16.7B vision-language GUI agent for lower-latency action generation.
  35. ExecCritic: Learn to Test, Test to Improve for Coding Agents
    ArXiv · AI/CL/LG ·
    ExecCritic separates test generation from repair and trains coding agents with role-specific RL inside a fail-closed scaffold.
  36. Training-Free Task Vectors for LLM Behavioral Control
    ArXiv · AI/CL/LG ·
    Training-Free Task Vectors map activation steering into rank-one weight edits without fine-tuning.
  37. Anthropic says Claude worked "largely autonomously" over 11 days to formalize the proof of Fermat's Last Theorem in the Lean programming language (Anthropic)
    Techmeme ·
    Anthropic reports Claude autonomously completed a Lean formalization of Fermat's Last Theorem, demonstrating long-horizon theorem-proving capability.
  38. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
    HF Daily Papers ·
    Qwen3.8-Flash-Next details a sparse 125B MoE architecture that cuts training compute while matching larger predecessors.
  39. Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
    HN · Frontpage AI ·
    Qwen introduced a Flash-Next architecture focused on lower-cost inference, attracting major developer discussion around model efficiency.
  40. LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
    HF Daily Papers ·
    LAION-BVD releases a 10-million-hour open video dataset aimed at large-scale multimodal pretraining across video, audio, and image tasks.
  41. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
    HF Daily Papers ·
    The Station multi-agent environment produced novel mathematical results across several AlphaEvolve-style construction problems.
  42. Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels (Terry Chen/NVIDIA Technical Blog)
    Techmeme ·
    Nvidia’s AVO agent system completed every ARC-AGI-3 public task, signaling a notable benchmark result for long-horizon autonomous coding agents.