[attvc keys ['entries', 'updatedAt', 'totalCount'] total=22]
{"entries": "list", "updatedAt": "2026-09-03T03:15:19.196Z", "totalCount": 22}
Pai 的 AI 雷达
今天:Gemini 3.8 Flash(Agent 原型),Muse Spark 1.3。并行还有 Qwen3.8-Max-0902 快照、Path to Astra、Claude Fable 5.1。
{"entries": "list", "updatedAt": "2026-09-03T03:15:19.196Z", "totalCount": 22}
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these
Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designe
RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual
When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no tru
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable m
Google AI 团队发布教程,讲解如何为 LLM-as-a-Judge 评测编写可靠的布尔式评分标准,指出模糊提示会导致评估不一致和浪费 token。文中给出四条经验:问题保持原子化且互不重叠、只让评判模型评估客观事实(可用 RFC 2119 术语如 MUST 表述)、只评 prompt 中明确要求的内容、用专家标注的 golden set 校准评判模型
Google DeepMind 发布 Gemini 3.8 Flash 与 3.8 Flash Cyber。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtkbdbti01n8roz5k5kt1g98
Meta 发布 Muse Spark 1.3,为五个月内第四个 Muse Spark 版本。max 变体(合作伙伴限时预览)在 Artificial Analysis Intelligence Index 得 62 分。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtkme19o02plro5q9
清华大学与美团学术论坛开放报名,主题对准物理世界 AI 与大模型,覆盖具身与世界模型。
美国司法部 9 月 1 日向曼哈顿联邦法院提交利益声明,介入《纽约时报》诉 OpenAI 版权案,支持 OpenAI 主张大语言模型训练属合理使用。DOJ 以国家安全和 AI 产业竞争力为主要论据,称训练使用具有转换性;《纽约时报》发言人批评政府站在 AI 公司一边牺牲创作者权益。法官要求双方最迟 9 月 4 日提交简易判决动议,该案结果可能为 AI 训练版
美国军方平台加上 ChatGPT 与 Grok。
问 LLM 能否创建并进化自己的 agent harness。
Verge:模型「更努力」,账单不一定更低。
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data s
按 TCO 估卡,别按峰值宣传买。
Assessing vulnerability detection tools for smart contracts requires datasets with known ground truth, yet such datasets are scarce and difficult to build by hand. We propose an ap
Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that con