AI论文速递 2026年08月11日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-08-11
📌 重点关注¶
- Douyin Multimodal Embedding Model Technical Report | arXiv — 【重点关注】 Multimodal representation learning is a cornerstone of modern AI. By encoding... 💡 工业级多模态嵌入实战,两阶段训练兼顾效率精度,移动端搜索推荐可借鉴
- Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection | arXiv — 【重点关注】 The malicious use of generative artificial intelligence to create highly real... 💡 多Agent协作检测deepfake,突破单模型视角局限,Agent架构设计有新意
- Motif 3: Technical Report | arXiv — 【重点关注】 We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 3... 💡 314B MoE大模型,GDLA新注意力机制,13B激活参数推理友好
📋 其他值得关注¶
- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States | arXiv — Learning-based memory systems for self-evolving LLM agents face two tightly c...
- OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction | arXiv — Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabil...
- YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family | arXiv — Generic parameter-efficient fine-tuning (PEFT) methods transferred from langu...
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills | arXiv — Robot learning is splitting into two bets: policies that bake competence into...
- StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding | arXiv — Deploying autonomous multimodal agents in continuous, real-world environments...
- CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models | arXiv — Benchmarking video-language models has largely focused on short clips and sin...
- PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say | arXiv — LLM-based agents are rapidly advancing, autonomously invoking external tools ...