AI论文速递 2026年07月31日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-31
📌 重点关注¶
-
Beacon: Knowing When and How to Perform Agentic Visual Reasoning | arXiv — 【重点关注】 The fundamental goal of agentic visual reasoning is to improve the success ra... 💡 教Agent学会"何时看、怎么看",用最小的视觉步骤提升任务成功率,对Agent设计有直接启发
-
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms | arXiv — 【重点关注】 Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph... 💡 BM25在大规模场景反超向量检索,知识库RAG不必只迷信embedding,简单方案值得回归
-
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM | arXiv — 【重点关注】 Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A... 💡 抛弃LLM中心化流水线,在消费级GPU跑32Hz实时VLA,端侧AI落地的新思路
📋 其他值得关注¶
- VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System | arXiv — Text-to-video models have achieved remarkable visual quality, yet they still ...
- OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis | arXiv — Biomedical image analysis spans diverse modalities and tasks, yet real-world ...
- SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch | arXiv — LLM-based agents excel at software engineering tasks where an existing codeba...
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents | arXiv — Training large language models (LLMs) to act in long-horizon games is a promi...
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding | arXiv — Large language model (LLM) agents are increasingly expected to assist users i...
- Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems | arXiv — Modern multi-agent knowledge systems increasingly accumulate knowledge throug...
- SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution | arXiv — Large language model agents often encounter related yet distinct tasks that s...