周报

周报

2026-W31 周报精选:Inside our 353,000-person vi;Circles powers telco persona;Latest open artifacts (#23):

条目
24
覆盖天数
2
当前批次
2026-W31
#14

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited

推荐理由「From RLVR to RLSVR: Task Transfo…」偏强化学习,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#15

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement

推荐理由「N_0-VTLA: Scaling Vision-Tactile…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#17

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation

推荐理由「Fewer Clarifications, Better Cod…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#19

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large sc

推荐理由「N_0-TWAM: Scaling Tactile-Native…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#20

Scaling Properties of Text Conditioning in Visual Generation

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Sur

推荐理由「Scaling Properties of Text Condi…」偏生成模型,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#21

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulat

推荐理由「AISPA: User-Centric System Promp…」对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#22

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but ca

推荐理由「SAF-OPD: Stable Advantage Fusion…」偏强化学习,对造物者有用的是方法能不能迁到自己的 Agent 或生成管线,不是把标题再念一遍。
研究与学习HF daily papers
#24

用850个单词统治世界:一个语言学家的疯狂实验

1930年,一位英国语言学家出版了一本书,声称用850个英语单词就能完成日常生活的全部表达。 这听起来不可思议。 英语词汇量超过17万,莎士比亚一个人就用了两万多个不同单词。 850个,够干什么? 但这个想法在二战结束后引发了全球范围的讨论,影响了BBC的广播方式,塑造了今天全球英语教学的基础词汇体系,甚至间接启发了乔

推荐理由「用850个单词统治世界:一个语言学家的疯狂实验」是硬件向。看开源、固件和动手步骤,发布稿不构成推荐。
硬件与机器人乔木博客