AI论文速递 2026年07月28日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-28
📌 重点关注¶
- Data Pyramid for Embodied Manipulation | arXiv — 【重点关注】 Multimodal foundation models learned to see and to speak by consuming the who... 💡 机器人数据金字塔:端侧Agent感知-行动闭环的数据策略参考
- ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding | arXiv — 【重点关注】 Multimodal large language models (MLLMs) hold immense potential to revolution... 💡 视觉多模态医疗LLM,融合架构可启发移动端AI诊断方向
- Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making | arXiv — 【重点关注】 Large language models are increasingly deployed as agents, but reliable agent... 💡 冻结LLM外挂控制头做Agent决策路由,启发性强,直接可用于你的框架设计
📋 其他值得关注¶
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking | arXiv — Embodied visual tracking (EVT) requires a mobile agent to continuously follow...
- Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills | arXiv — LLM training is shifting from manual design and annotation to interaction-dri...
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents | arXiv — Traditional agent development is split across prompt templates, tool schemas,...
- OpenForgeRL: Train Harness-native Agents in Any Environment | arXiv — Modern AI agents rely on elaborate inference harnesses such as Claude Code, C...
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search | arXiv — Agentic search enables large language models to solve knowledge-intensive tas...
- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems | arXiv — Production AI agents' failures are less often due to an inability to reason w...
- Multi-Turn On-Policy Distillation with Prefix Replay | arXiv — We study on-policy distillation (OPD) for agentic tasks, where an LLM agent i...