AI论文速递 2026年07月29日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-29
📌 重点关注¶
- Data Pyramid for Embodied Manipulation | arXiv — 【重点关注】 Multimodal foundation models learned to see and to speak by consuming the who... 💡 金字塔式数据策略,从互联网多源数据到真实机器人数据层层递进,为具身智能构建可扩展训练管线
- ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding | arXiv — 【重点关注】 Multimodal large language models (MLLMs) hold immense potential to revolution... 💡 以视觉为中心的医学多模态LLM,融合多种影像模态做端到端诊断,展示垂直领域MLLM落地方案
- Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making | arXiv — 【重点关注】 Large language models are increasingly deployed as agents, but reliable agent... 💡 冻结LLM上加轻量控制头,从隐藏态直接推断工具调用/澄清/拒绝决策,对Agent架构设计极有启发
📋 其他值得关注¶
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models | arXiv — Large Vision-Language Models (LVLMs) remain bottlenecked by massive computati...
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search | arXiv — Relevance is a query-dependent estimate of whether a document or excerpt cont...
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search | arXiv — Agentic search enables large language models to solve knowledge-intensive tas...
- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems | arXiv — Production AI agents' failures are less often due to an inability to reason w...
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model | arXiv — Standard vision-language models (VLMs) suffer from Moravec's paradox: they ex...
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents | arXiv — The alignment of Small Language Models (SLMs) in the 70--500M parameter range...
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models | arXiv — We introduce PerceptionBench, a benchmark specifically designed to evaluate t...