Uncut

Physics remains a difficult testbed for multimodal models because it mixes diagrams, text, equations, and multi-step reasoning. OmniPhys adds a large Chinese…

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

#4 · benchmark · Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang ·

A 15,246-question multimodal physics benchmark built from Chinese educational corpora across middle school to university levels.

Physics remains a difficult testbed for multimodal models because it mixes diagrams, text, equations, and multi-step reasoning. OmniPhys adds a large Chinese educational benchmark, which could expose gaps that English-heavy and general visual reasoning sets miss.

new arXiv paper with repository link

uncovered in mainstream sources

Recent Uncut picks

  1. Prefix Sliding for efficient test-time scaling
    #8 · paper · A test-time scaling method that drops older reasoning-prefix tokens to reduce long-reasoning attention cost.
  2. SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
    #9 · benchmark · A benchmark for testing how multimodal LLMs follow complex instructions across scientific tasks and disciplines.
  3. ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives
    #10 · paper · A dual-agent framework for evidence-aware question answering over long literary and narrative texts.
  4. Learning New Facts with QLoRA: An Acquisition-Retention Frontier
    #11 · paper · A controlled study of how QLoRA rank affects learning new facts while retaining unrelated capabilities.
  5. Controlling for Omitted Variable Bias in Deep Neural Networks
    #12 · paper · A deep-learning method for controlling known confounders and omitted-variable bias in neural predictions.
  6. GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
    #3 · paper · An inference-time method for reducing demographic bias in generative VLMs by steering along a counterfactual bias subspace.
  7. OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
    #4 · benchmark · A 15,246-question multimodal physics benchmark built from Chinese educational corpora across middle school to university levels.
  8. Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness
    #5 · paper · A machine-unlearning paper arguing that update alignment predicts whether forgotten LLM knowledge can be relearned.
  9. MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
    #6 · benchmark · A diagnostic Lean 4 theorem-proving benchmark spanning 13 undergraduate and graduate math domains.
  10. CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
    #7 · paper · A counterfactual-causal skill graph approach for retrieving reusable agent skills without prompting the full library.
  11. ByteBunny777/Docguard
    #2 · repo · A Go CLI that scans documents for hidden prompt-injection payloads before they enter a RAG pipeline.
  12. thomasbek3/hermes-computer-viewer
    #1 · repo · A Hermes Desktop plugin that adds a live KVM-style remote desktop pane across cloud and LAN machines.
  13. kuliantnt/qq-maid-bot
    #42 · repo · A local Rust service for a general-purpose QQ bot, tagged around OneBot11, low-memory use, RAG, and wxbot.
  14. noyce6983981-max/ocr-vlm-local-retrieval
    #40 · repo · A local-first prototype for OCR plus vision-language document retrieval, with an independent evaluation angle.
  15. avar6/GLM-5.3-Flash-BF16-gguf
    #41 · model · A GGUF BF16 quantization of GLM-5.3-Flash aimed at conversational and endpoint-compatible use.
  16. blackhaiyu-sudo/specrag
    #39 · repo · A Python evidence-based RAG knowledge base for PRDs, business rules, SOPs, process docs, and product screenshots.
  17. raiyanyahya/llmaker
    #38 · repo · A Go CLI for self-hosting a modern LLM stack from the terminal.
  18. goobolabs/somali-language-standard
    #37 · benchmark · A versioned, machine-readable Somali language standard spanning orthography, grammar, terminology, translation, and AI resources.
  19. K-Dense-AI/scientific-agents
    #36 · repo · A set of AGENTS.md profiles that encode expert scientific and engineering reasoning styles for AI agents.
  20. Agent-Field/reels-af
    #35 · repo · A Python multi-agent system for automating short-form video creation at a claimed low per-reel cost.
  21. YintongHuo/awesome-agent-trajectory
    #33 · benchmark · A curated collection of agent trajectory analysis techniques and benchmarks for studying how LLM agents behave over time.
  22. Krypto-Whitehat/qwen3.8-9b-uncensored-cyber-exploit-XRPL-v3
    #34 · model · A Qwen-derived text-generation model packaged for cybersecurity and XRPL bug-triage workflows.
  23. yuwen-cool/ywcrew
    #31 · repo · A TypeScript orchestration tool that dispatches tasks to locally subscribed AI coding agents in parallel.
  24. Baekpica/Qwen3.8-Flash-Next-Mixed-Quant-SSD-PLE-GGUF
    #32 · model · A GGUF mixed-quant version of Qwen3.8-Flash-Next for image-text-to-text use, tagged for SSD offload.
  25. Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis
    #26 · paper · Adapts a pretrained latent diffusion model to synthesize cardiac MRI conditioned on clinical metadata and slice position.
  26. Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets
    #27 · paper · Tests whether updating only 9% of parameters can adapt EEG foundation models under fixed clinical compute budgets.
  27. Show HN: Mole – Deep research agent for your terminal
    #28 · repo · A terminal-based deep research agent that drew meaningful Hacker News discussion as a Show HN launch.
  28. Hilbert-beinghappy/seektty
    #29 · repo · A JavaScript terminal UI for DeepSeek Harness, aimed at making DeepSeek-based coding-agent workflows pluggable from the CLI.
  29. AliAkrami1375/Li-Translate
    #30 · repo · A Vue-based platform for AI subtitle generation and natural-language translation for video and audio.
  30. dondai44423/donsetch
    #21 · repo · A Rust web fetch, search, and crawl tool for AI agents that avoids API keys and external accounts.