AI论文速递 2026年07月30日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-30
📌 重点关注¶
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models | arXiv — 【重点关注】 Large Vision-Language Models (LVLMs) remain bottlenecked by massive computati... 💡 首个直接针对端侧延迟优化的LVLM视觉编码器,金字塔架构+异构空间混合器,移动端AI部署利器
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM | arXiv — 【重点关注】 Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A... 💡 将V→L→A路径重构为V+L→A直接映射,4090上32Hz实时推理,机器人端侧轻量部署新范式
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search | arXiv — 【重点关注】 Relevance is a query-dependent estimate of whether a document or excerpt cont... 💡 将相关性从检索排序升级为语料交互执行先验,让Agent更聪明地搜索,RAG系统可直接受益
📋 其他值得关注¶
- OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis | arXiv — Biomedical image analysis spans diverse modalities and tasks, yet real-world ...
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model | arXiv — Standard vision-language models (VLMs) suffer from Moravec's paradox: they ex...
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents | arXiv — Training large language models (LLMs) to act in long-horizon games is a promi...
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding | arXiv — Large language model (LLM) agents are increasingly expected to assist users i...
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents | arXiv — The alignment of Small Language Models (SLMs) in the 70--500M parameter range...
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models | arXiv — We introduce PerceptionBench, a benchmark specifically designed to evaluate t...
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs | arXiv — Parametric retrieval enables LLMs to retrieve tools implicitly by assigning e...