AI论文速递 2026年08月08日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-08-08
📌 重点关注¶
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models | arXiv — 【重点关注】 Computer-using agents (CUAs) are advancing rapidly across the digital world. ... 💡 建立跨平台CUA奖励模型标准化评估体系,Agent评测必备参考
- SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding | arXiv — 【重点关注】 Understanding 3D scenes is fundamental to embodied intelligence, requiring jo... 💡 动态多模态编排的3D场景理解,端侧AI空间感知参考方向
- K-EXAONE 2.0 Technical Report | arXiv — 【重点关注】 This technical report presents K-EXAONE 2.0, an open-weight multilingual foun... 💡 韩国开源多语言基础模型技术报告,模型选型与微调参考
📋 其他值得关注¶
- From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models | arXiv — Economic World Models (EWMs) are generative economic models that simulate how...
- Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval | arXiv — Unified multimodal retrieval aims to identify candidates that satisfy complex...
- Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains | arXiv — Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major...
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills | arXiv — Robot learning is splitting into two bets: policies that bake competence into...
- ChronoVision: Temporal Reasoning via Latent State Reconstruction | arXiv — Multimodal large language models excel at passive perception but struggle wit...
- FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory | arXiv — GUI agents must remember both useful experience from earlier tasks and unfini...
- BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation | arXiv — Leveraging pre-trained vision-language models (VLMs) to construct vision-lang...