Skip to content

AI论文速递 2026年07月10日(HuggingFace Daily Papers)

数据来源:https://huggingface.co/papers 采集时间:2026-07-10

📌 重点关注

  1. MentalThink: Shaping Thoughts in Mental SVG World | arXiv【重点关注】 We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Mu... 💡 符号空间推理新范式,Agent内部思维可视化值得借鉴
  2. Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE | arXiv【重点关注】 Modern LLMs are increasingly deployed in long-context applications such as re... 💡 动态RoPE低成本扩展长上下文,Agent长记忆场景刚需
  3. HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better | arXiv【重点关注】 We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-la... 💡 腾讯轻量OCR-VLM,端侧AI应用部署好参考

📋 其他值得关注

  1. Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing | arXiv — Academic output is produced across a fragmented toolchain: literature discove...
  2. CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration | arXiv — Complex image creation and editing often require more than a single generatio...
  3. RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures | arXiv — Pretrained video generative models are promising backbones for visuomotor con...
  4. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation | arXiv — We present AgentLens, a production-assessed benchmark for interactive code ag...
  5. Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs | arXiv — Touch supplies the physical grounding needed to perceive intrinsic material p...
  6. When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers | arXiv — LLM agents increasingly rely on retrieval buffers to store and reuse past exp...
  7. Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation | arXiv — Mainstream Vision-Language-Action (VLA) models predict actions primarily from...