AI论文速递 2026年07月10日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-10
📌 重点关注¶
- MentalThink: Shaping Thoughts in Mental SVG World | arXiv — 【重点关注】 We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Mu... 💡 符号空间推理新范式,Agent内部思维可视化值得借鉴
- Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE | arXiv — 【重点关注】 Modern LLMs are increasingly deployed in long-context applications such as re... 💡 动态RoPE低成本扩展长上下文,Agent长记忆场景刚需
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better | arXiv — 【重点关注】 We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-la... 💡 腾讯轻量OCR-VLM,端侧AI应用部署好参考
📋 其他值得关注¶
- Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing | arXiv — Academic output is produced across a fragmented toolchain: literature discove...
- CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration | arXiv — Complex image creation and editing often require more than a single generatio...
- RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures | arXiv — Pretrained video generative models are promising backbones for visuomotor con...
- AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation | arXiv — We present AgentLens, a production-assessed benchmark for interactive code ag...
- Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs | arXiv — Touch supplies the physical grounding needed to perceive intrinsic material p...
- When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers | arXiv — LLM agents increasingly rely on retrieval buffers to store and reuse past exp...
- Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation | arXiv — Mainstream Vision-Language-Action (VLA) models predict actions primarily from...