AI论文速递 2026年08月04日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-08-04
📌 重点关注¶
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger | arXiv — 【重点关注】 Multimodal agents for visual question answering increasingly operate as multi...
💡 证据溯源约束Agent推理链,对构建可信AI Agent极具参考价值
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement | arXiv — 【重点关注】 Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progr...
💡 自验证奖励让LLM在开放任务中自我进化,RL新范式值得跟踪
- Constitutional Midtraining: Content Presence Drives Alignment Gains | arXiv — 【重点关注】 Post-training alignment is often shallow, eroding under fine-tuning. Whether ...
💡 中训练阶段注入对齐优于传统后训练,对齐工程新思路
📋 其他值得关注¶
- Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions | arXiv — Deep Research agents extend LLM-based assistants into long-horizon workflows ...
- EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents | arXiv — The web is increasingly accessed by AI agents rather than humans. Every agent...
- Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability | arXiv — Deployed LLM agents increasingly keep their long-term memory as a filesystem:...
- SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing | arXiv — Autonomous multi-vehicle racing requires real-time planning of diverse compet...
- See2Think: Do Multimodal Models Really Use Intermediate Visual States? | arXiv — Multimodal large language models increasingly use sketches, annotations, tool...
- ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step | arXiv — To operate robustly in open-world environments, autonomous agents should be a...
- Weak-to-Strong On-Policy Distillation | arXiv — On-policy distillation (OPD), which aligns a student with the teacher's token...