周报

周报

2026-W32 周报精选:Fay:fay是一个帮助数字人(2.5d、3d、移动、p;How Zapier transformed core ;Meta is back with Muse Glimm

条目
80
覆盖天数
7
当前批次
2026-W32
#35

CEO上个月在 AI 上花了 1.3 万美元,然后说「自动化是个谎言」

Dan Shipper 上个月的 Codex 账单是 13,000 美元,他的 COO 给了他一个白眼。 但他不打算减少用量。 Every 的 Slack 里现在有 20 个 agent 和 20 个人类,一起工作。 代码、写作、客服、邮件、研究备忘录,能用 AI 处理的流程全部交给了 agent。 27 个全职员工,

推荐理由「CEO上个月在 AI 上花了 1.3 万美元,然后说「自动化是个…」可当助手或自动化原型来拆,看交互和工具边界,不是发布会。
创作与自媒体乔木博客
#40

MIT博士“叛逃“物理系:我要用物理学重写AI的底层逻辑

https://www.xiaoyuzhoufm.com/episode/6a69b07eb581962ce2bd4d97 2024 年 4 月,一篇挂在 arXiv 上的论文在技术社群里炸开了锅。 有人说,统治深度学习几十年的 MLP(多层感知机)可能要被改写了。 也有人说,这不过是又一个"看起来很美"的架构。 这篇

推荐理由「MIT博士“叛逃“物理系:我要用物理学重写AI的底层逻辑」里能借鉴的是做法和约束,不是又一条行业新闻。
创作与自媒体乔木博客
#46

vLLM 背后的人:一个开源项目如何长成一家公司

这期商业访谈录也超级精彩,了解到很多关于 Deepseek 的厉害之处。 小宇宙博客地址: https://www.xiaoyuzhoufm.com/episode/6a66ed17a3fec224d5a3f744 AI 重写成文章,帮没空听的朋友节省3小时: 有人在公司注册前夕接到电话,对方开门见山:你们四个创始人,

推荐理由「vLLM 背后的人:一个开源项目如何长成一家公司」看能不能进现有工作流(能力、接口、价格),不当通稿收藏。
创作与自媒体乔木博客
#48

用850个单词统治世界:一个语言学家的疯狂实验

1930年,一位英国语言学家出版了一本书,声称用850个英语单词就能完成日常生活的全部表达。 这听起来不可思议。 英语词汇量超过17万,莎士比亚一个人就用了两万多个不同单词。 850个,够干什么? 但这个想法在二战结束后引发了全球范围的讨论,影响了BBC的广播方式,塑造了今天全球英语教学的基础词汇体系,甚至间接启发了乔

推荐理由「用850个单词统治世界:一个语言学家的疯狂实验」是硬件向。看开源、固件和动手步骤,发布稿不构成推荐。
硬件与机器人乔木博客
#52

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple

推荐理由「SimWAM: A Simple World Action Mo…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#57

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement with CausalNav, a controller built around a

推荐理由「CausalNav: Reliability-Certified…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#58

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evi

推荐理由「CliniCARE-Bench: Clinical Calibr…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#59

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate recommendations conditioned on item-interaction histories. In these sys

推荐理由「Preserving Item Semantics for Fr…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#63

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge

推荐理由「When the Judge Should Not Decide…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#65

CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals

Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathing, posture, sensor contact, and true clinical deterioration. This paper asks when

推荐理由「CFD-Guided Detection of Concept…」偏多模态,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#66

Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. Understanding th

推荐理由「Who Built This Model? Tracing LL…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#67

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic task

推荐理由「AgentOPSD: Recursive Self-Distil…」偏强化学习,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#68

The Anatomy of a Prompt Injection: A Component Model for Structured Analysis

Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embed

推荐理由「The Anatomy of a Prompt Injectio…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#69

WorldClaw: Agentic 3D Open-World Generation at Scale

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstrea

推荐理由「WorldClaw: Agentic 3D Open-World…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#70

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits report a binary out-of-domain rate; none ask whether the model knew. We jointly audit hallucination rate (OOD@10)

推荐理由「Do LLM Recommenders Know When Th…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#71

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA eval

推荐理由「OSReward: Instituting Standardiz…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#72

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights

推荐理由「Interpretable MEG Decoding of Pe…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#73

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic o

推荐理由「Who Verifies the Benchmark? Dece…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#74

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators

推荐理由「EnvACE: Internalizing Environmen…」偏强化学习,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#75

The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era

Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a curriculum-framework paper, grounded in a structured narrative review of labor-mark

推荐理由「The Capability Ladder: A Curricu…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习arXiv
#77

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them

推荐理由「HarnessOpt-Bench: Evaluating LLM…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#78

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through

推荐理由「From Economic Agents to Agentic…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#79

The Utility of LLMs in Recommender Systems Explanation Evaluation

Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best f

推荐理由「The Utility of LLMs in Recommend…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习papers.cool cs.AI Atom
#80

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual s

推荐理由「GST-Bench: Can VLMs Develop Glob…」偏Agent,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers