EarlyEval:提前预测 Agent 评测结果
2609.02783:完整跑一遍前沿 Agent 基准太贵,用早期轨迹预测最终成败来省钱。
AI 资讯
分类页是滚动库存,每页 24 条。首页只放当天精选。
2609.02783:完整跑一遍前沿 Agent 基准太贵,用早期轨迹预测最终成败来省钱。
2609.02886,今日 HF daily 最高票(32)。开源数据和可扩展训练配方。
2609.02737:模型实际只盯上下文一小截,却仍扫整份 KV。让模型决定听哪些 token。
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions.
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We
用 ASCII 艺术重包装有害请求。
内生授权清洗:记忆层变成越权面。
按结构质量做输入自适应稀疏 prefilling。
用首次失误学过程奖励,不必等结局。
分层错误记忆,按诊断让模型自进化。
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce D
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Proc
短剧生成的端到端评测。
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an
按事件组织多模态记忆。
测代码 embedding 检索和功能正确之间的缝。
评 LLM 危险能力的框架。
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they ar
一次前向做细粒度 caption,配 SimLoss。
把语言理解接到世界控制。
连续多日软件开发,套件还能自我改进。
问 LLM 能否创建并进化自己的 agent harness。
从 in-the-wild 视频生成动物运动。
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to