Skip to content

AI论文速递 2026年08月04日(HuggingFace Daily Papers)

数据来源:https://huggingface.co/papers 采集时间:2026-08-04

📌 重点关注

  1. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger | arXiv【重点关注】 Multimodal agents for visual question answering increasingly operate as multi...

    💡 证据溯源约束Agent推理链,对构建可信AI Agent极具参考价值

  2. From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement | arXiv【重点关注】 Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progr...

    💡 自验证奖励让LLM在开放任务中自我进化,RL新范式值得跟踪

  3. Constitutional Midtraining: Content Presence Drives Alignment Gains | arXiv【重点关注】 Post-training alignment is often shallow, eroding under fine-tuning. Whether ...

    💡 中训练阶段注入对齐优于传统后训练,对齐工程新思路

📋 其他值得关注

  1. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions | arXiv — Deep Research agents extend LLM-based assistants into long-horizon workflows ...
  2. EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents | arXiv — The web is increasingly accessed by AI agents rather than humans. Every agent...
  3. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability | arXiv — Deployed LLM agents increasingly keep their long-term memory as a filesystem:...
  4. SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing | arXiv — Autonomous multi-vehicle racing requires real-time planning of diverse compet...
  5. See2Think: Do Multimodal Models Really Use Intermediate Visual States? | arXiv — Multimodal large language models increasingly use sketches, annotations, tool...
  6. ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step | arXiv — To operate robustly in open-world environments, autonomous agents should be a...
  7. Weak-to-Strong On-Policy Distillation | arXiv — On-policy distillation (OPD), which aligns a student with the teacher's token...