AI论文速递 2026年08月06日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-08-06
📌 重点关注¶
- Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent | arXiv — 【重点关注】 We introduce Video-DeepResearch (Video-DR), extending multimodal agents from ... 💡 多模态DeepResearch从文本扩展到视频理解,Agent能力边界持续突破
- TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning | arXiv — 【重点关注】 Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through i... 💡 后见之明优化长程工具调用训练,为Agent推理提供更精细的奖励信号
- ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation | arXiv — 【重点关注】 Text-to-image (T2I) models can produce visually compelling images, yet they r... 💡 推理+工具+生图统一策略,Agent范式在图像生成领域的完整落地
📋 其他值得关注¶
- Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements | arXiv — Do Large Language Models (LLMs) possess genuine structural reasoning, or mere...
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation | arXiv — MLLM-based segmentation faces a core segmentation trilemma: high segmentation...
- BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation | arXiv — Leveraging pre-trained vision-language models (VLMs) to construct vision-lang...
- OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents | arXiv — LLM agents are increasingly applied to open-ended everyday requests that span...
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents | arXiv — Recursive self-improvement requires agents to turn accumulated experience int...
- CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning | arXiv — Chart question answering (CQA) requires multimodal large language models (MLL...
- PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs | arXiv — Scientific poster construction compresses a long multimodal paper into a read...