Megadose AI progress, ranked daily.

Uncut

Most AI-content detection discussion centers on language cues, which are brittle once text is edited or paraphrased. This work shifts the signal to web…

Show HN: Training a model to identify AI web content from structure alone

#1 · paper · jochenmadler ·

A paper exploring whether AI-generated web content can be detected from page structure rather than text alone.

Most AI-content detection discussion centers on language cues, which are brittle once text is edited or paraphrased. This work shifts the signal to web structure, a different layer that may capture production patterns rather than prose style. If the approach holds up, it could matter for spam analysis, provenance research, and measuring how AI-generated pages spread across the web.

HN front page · 74 points and 28 comments

1 mainstream mention

Recent Uncut picks

  1. TencentCloud/octop-memory
    #8 · repo · A Python toolkit for giving agents persistent memory: preferences, session context, and portable memory across agent systems.
  2. Talorys – A self-hosted personal AI agent on Cloudflare's free tier
    #7 · repo · A self-hosted personal AI agent designed to run on Cloudflare's free tier.
  3. dongnguyenvie/BashCut
    #4 · repo · A native macOS video editor designed so coding agents can control timeline editing, captions, voiceover, and plugins through CLI or MCP.
  4. Akun-python/autumn-recruitment
    #5 · repo · A Chinese study repo tying algorithms, ML, deep learning, LLMs, agents, AI infra, speech, and data science into runnable notes.
  5. kkjiangk/agent-py
    #6 · repo · A personal AIOps project combining durable workflows, hybrid retrieval, event delivery, and verified API/browser contracts.
  6. Show HN: Training a model to identify AI web content from structure alone
    #1 · paper · A paper exploring whether AI-generated web content can be detected from page structure rather than text alone.
  7. Show HN: Apogee: Rebuilding Mozilla's Orbit, fully local and private
    #2 · repo · A local, private rebuild of Mozilla Orbit-style browsing assistance packaged as an open GitHub project.
  8. Show HN: Nightwatch – a Mac menu-bar app that tells you when tonight is clear
    #3 · repo · A macOS menu-bar app that turns local sky conditions into a quick signal for whether tonight is clear.
  9. HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
    #26 · paper · Proposes higher-order tensor modeling to preserve multimodal synergy that cannot be recovered from single modalities alone.
  10. wh000wh000/awesome-jev-live
    #27 · repo · Maintains an evidence-graded, frequently rebuilt index of TypeSafe System One SDKs, MCP tools, agents, apps, and open models.
  11. Router0824/OpsPilot-AI
    #28 · repo · Packages meetings, knowledge, decisions, tasks, risks, workflows, and org memory into an open-source AI operations workspace.
  12. solintellegence/sol-pro
    #29 · model · Releases a PyTorch text-generation small language model using recurrent depth and grouped-query attention.
  13. BinceQu/RoboHarness
    #21 · repo · Gives LLM agents a visual-geometric interface for understanding and invoking embodied robotic tasks.
  14. Let your AI agents paint big arrows, boxes and text on your screen
    #22 · repo · Lets agents draw arrows, boxes, and text directly on a user's screen to guide attention or explain actions.
  15. sno-ai/sno-station-skills
    #23 · repo · Packages 31 tested skills for coordinating Claude Code and Codex across review, handoff, briefs, and work orders.
  16. kouhxp/gutsy
    #24 · repo · Runs a local 0.8B GGUF decision model that returns calibrated probabilities for yes/no, choice, and scoring questions.
  17. Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows
    #25 · paper · Proposes a unified forecasting backbone for numerical multivariate signals and text-assisted narrative-flow time series.
  18. mhtsec/ARTEX
    #16 · repo · An autonomous AI penetration-testing system from Baidu's agent+ attack-defense challenge, implemented in Go.
  19. FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
    #17 · benchmark · A benchmark for testing whether streaming video-language models can perceive fast real-world events under context limits.
  20. LosaLosSantos/aurelio-finance
    #18 · repo · A local-first personal finance app that adds an AI advisor on top of net worth, investments, ETFs, cash, and debt tracking.
  21. Show HN: Durable Actors – OSS Durable Objects with configurable compute
    #19 · repo · An open-source take on Durable Objects-style actors with configurable compute.
  22. Prism-Shadow/travel-agent
    #20 · repo · A cross-platform desktop travel agent that uses browser automation to plan trips and stop before payment.
  23. One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
    #11 · paper · An attack study showing how caller-defined option text can break typed decision models used as agent guardrails.
  24. Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology
    #12 · benchmark · A theory-guided benchmark for probing how LLMs assign social cause, responsibility, blame, and credit.
  25. MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation
    #13 · paper · A spectra-conditioned molecular structure method that augments diffusion generation with motifs and search-tree inference.
  26. ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
    #14 · paper · A VLM-agent training method that learns reusable visual-native skills instead of reducing spatial strategy to text.
  27. Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
    #15 · paper · A dense correspondence framework for identity-preserving matches when editing or generation breaks physical continuity.
  28. GRPODropout: Less is More for Online Reinforcement Learning Rollouts
    #6 · paper · A GRPO variant that tries to reduce entropy collapse in online RL rollouts by changing rollout behavior rather than only the loss.
  29. DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
    #7 · benchmark · A benchmark for testing whether multimodal LLMs can understand animated chart videos and their narrative structure.
  30. Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits
    #8 · paper · A systematic study of test-time compute strategies for tabular foundation models, including a tiny trainable adaptation method.