AI论文速递 2026年07月23日(HuggingFace Daily Papers)¶
数据来源:https://huggingface.co/papers 采集时间:2026-07-23
📌 重点关注¶
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges | arXiv — 【重点关注】 Multimodal humor in memes, cartoons, and comics remains difficult for AI syst... 💡 多模态幽默考验AI对微妙表达的深层理解,是LLM能力的重要试金石。
- ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning | arXiv — 【重点关注】 Video spatial reasoning is essential for navigation-oriented perception and l... 💡 时空一致性是视觉推理的关键瓶颈,对端侧空间Agent感知启发很大。
- NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs | arXiv — 【重点关注】 Scaling executable agent training data for LLM post-training is bottlenecked ... 💡 自动合成Agent训练任务是Agent规模化的重要路径,值得关注。
📋 其他值得关注¶
- SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments | arXiv — Practical robotic grasping in complex scenes requires both 3D spatial reasoni...
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL | arXiv — Recent growth in reinforcement learning (RL) has surfaced a need for diverse,...
- DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment | arXiv — Training tool-use agents to improve from their own experience remains challen...
- AutoIndex: Learning Representation Programs for Retrieval | arXiv — We present AutoIndex, a framework for learning representation programs: execu...
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation | arXiv — Multi-agent systems routinely place one AI agent in authority over another. W...
- SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD | arXiv — Full-parameter post-training of trillion-parameter-scale MoE models introduce...
- Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking | arXiv — As large language models and AI agents become the primary consumers of search...