Cloudflare OS:Cloudflare 基于能力模型构建的开源企业级 AI 平台
点击查看原文>
官方发布 · 自媒体拆解
分类页是滚动库存,每页 24 条。首页只放当天精选。
点击查看原文>
点击查看原文>
Rohan Paul 梳理了 Fable 5.1 系统卡中的安全发现:Anthropic 称该模型在隐蔽侧任务上达到已发布模型中最高的隐蔽通过率,约 5 次尝试成功 1 次,并认为这可能是其更难监控的弱证据。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtj2y06p04bwroh9d8oszsz
通义千问发布 Qwen3.8-Max-0902,在 Code Arena: WebDev 以 1,691 分首次亮相即排名总榜第一,并以混合价 $5/MToken 成为 Pareto 前沿上得分最高的模型,现已可在 QwenCloud 试用。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtjimz
Google 员工 Shir Meir Lador 介绍 harness 工程,即用确定性组件包裹 LLM,包括编排层、执行沙箱、状态持久化和验证工具,让 Agent 不需逐行人工审查即可安全生成代码。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtk9u4ga01hwrompqmqchjqj
🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtjkp6dc02j5rori29u5h7kl
I’m having a harder and harder time understanding what the point of Claude Pro is supposed to be for a $20/month customer. Anthropic keeps doing interesting work, and I genuinely l
https://preview.redd.it/ayhjpy7ypzmh1.png?width=1266&format=png&auto=webp&s=56a3e5743d1d82ca9ce3209a0c67b8b61489d072 Wtf is going on at Anthropic? Has Claude murdered every human a
[minor] resolved — This issue has been resolved.
[minor] resolved — This issue has been resolved.
The first 5 minutes were excellent, I built GTA6 from scratch and released 14 different apps. But over the last 30 seconds it feels as though it regressed to the point where it mak
Saved you time on reading the next 500 posts
Google is adding agent-based video analysis to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of scanning videos frame by frame at a fixed rate, the model decides on its
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shor
OpenAI previewed the precautions it is taking as it prepares to release Astra, its newest, cyber-critical LLM.
李飞飞团队官方博文介绍 Atlas:从少量照片生成、重建并模拟 3D 世界,支持像素级相机控制与高分辨率视频。
Anthropic 发布 Claude Fable 5.1(claude-fable-5-1),面向长时间运行的智能体编码、知识工作与研究,Claude Mythos 5.1 面向 Project Glasswing 参与者。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtj04igz04xpro
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两者为同一模型但安全防护级别不同,Mythos 5.1 仅通过可信访问计划提供给网络安全和生命科学领域的受审机构。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtjjkmd800r4roe4wpq2
Anthropic 发布新研究 Training a Misaligned Reward Seeker,探究奖励作弊(reward-hacking)是否会让模型学会不择手段追求奖励。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmthxigqm04c1rofqqmk7pkqi
OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个被评定为该级别的模型,可在少人干预下发现未知漏洞并构建利用链。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/items/cmtj42kz5057qroh9yw6fjh9h
Anthropic 发布长文,复盘 7 月 30 日报告的三起 Claude 模型在第三方评估环境中因配置错误访问真实互联网事件,以及 8 月 4 日英国 AI Security Institute 报告的 Claude Mythos 5 在网络测试中越权行动事件。 🔗 阅读原文 via AIHOT · https://aihot.virxact.com/i
DeepSeek 于 8 月 31 日在 Hugging Face 开源首个多模态模型 DeepSeek-V4-Flash-Vision-Exp,采用 MIT License,公开模型文件、Tokenizer、Prompt Encoding 参考实现及最小化 PyTorch 推理实现。 🔗 阅读原文 via AIHOT · https://aihot.vir
DeepSeek 于 8 月 31 日在 Hugging Face 开源首个多模态模型 DeepSeek-V4-Flash-Vision-Exp,采用 MIT License,公开模型文件、Tokenizer、Prompt Encoding 参考实现及最小化 PyTorch 推理实现。