🤖 AI Agent 研究Research
座席的使用正在转向编码和科学发现等长远任务,其中终端任务尤为重要。我们引入了T1 ,这是一种专家混合模型,总数为122B ,经过强化学习训练,在云沙箱中操作真实shell ,每个任务最多可调用300多个工具,通过执行每个任务自己的验证器获得奖励。
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier.
自动化实证研究是人工智能的长期发展方向。最近的自动研究( AutoResearch )代理实现了这一目标,因为现代LLM展示了独立实施解决方案并从执行结果中学习的能力。
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes.
我们研究LLM代理中的策划,其中代理隐蔽地追求不一致的目标。我们的重点是了解策划是如何从关键因素的相互作用中产生的,例如工具目标、环境承受能力、监督条件和可感知的后果。
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences.
顺序存储代理通过一个接一个地读取块来处理长文档,同时保持紧凑的内存状态,将文档遍历耦合到推理深度。这种耦合引入了对证据放置的敏感性,并将推理延迟与文档长度线性联系起来。
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length.
⭐ GitHub 热门项目GitHub Trending
【GitHub】SimonAgent —用纯Python从头开始构建的轻量级、安全第一的AI Agent运行时。具有声明式工具编排、双层上下文压缩、持久内存和按需技能系统。零框架,零魔法(⭐ 11 )
【GitHub】SimonAgent — A lightweight, security-first AI Agent runtime built from scratch in pure Python. Features declarative tool orchestration, dual-layer context compression, persistent memory, and an on-demand skill system. Zero frameworks, zero magic (⭐ 11)
【GitHub】AI智能体线束演进的精彩研究论文列表:基础、基准、配方、位置(⭐ 22 )
【GitHub】Awesome Research Paper List for AI agent harness evolution: Foundation, Benchmark, Recipe, Position (⭐ 22)
【GitHub】Claude代码技能:代码库审核、Hacker News摘要、GitHub趋势摘要、开源状态(⭐ 0 )
【GitHub】Claude Code skills: codebase audit, Hacker News digest, GitHub trending digest, and open-source status (⭐ 0)
【GitHub】使用AI代理( roblox-ts、Rojo、Studio MCP、Open Cloud、OpenAI、Claude-Code )构建Roblox游戏的开源桌面应用程序(⭐ 1 )
【GitHub】Open-source desktop app for building Roblox games with AI agents (roblox-ts, Rojo, Studio MCP, Open Cloud, OpenAI, Claude-Code) (⭐ 1)
【GitHub】MacOS的Claude Code & Codex状态栏。使用本机菜单栏控件和桌面小部件跟踪AI编码会话、使用限制和活动。免费开源。(⭐ 1 )
【GitHub】Claude Code & Codex status bar for macOS. Track AI coding sessions, usage limits and activity with native menu bar controls and desktop widgets. Free and open source. (⭐ 1)
🚀 模型与行业动态Models & Industry
首席执行官萨姆·奥尔特曼( Sam Altman )表示,虽然OpenAI已秘密申请IPO ,但该公司今年不会上市。
While OpenAI has filed confidentially for an IPO, the company will not be going public this year, according to CEO Sam Altman.
Anthropic的Dario Amodei和OpenAI的Sam Altman似乎都认为是时候“放慢前沿(AI)步伐”了。“那到底是什么样子?
Anthropic's Dario Amodei and OpenAI's Sam Altman seem to agree that it's time to "pace the frontier." What would that actually look like?
Meta的最新应用程序Muse的开始速度比该公司的其他应用程序(如Meta AI或Threads )慢。
Meta's newest app Muse is off to a slower start than the company's other apps, like Meta AI or Threads.
进入机器人的大脑,试图说服互联网这是人类。。Anthropic揭示了流氓AI代理和你一样讨厌CAPTCHA。
Come inside the mind of a bot trying to convince the internet it's human.. Anthropic reveals rogue AI agents hate CAPTCHAs, just like you.
🔥 社区热议Community
【Lobsters】热度: 15↑ | 14 评论 | 标签: vibecoding
【Lobsters】热度: 15↑ | 14 评论 | 标签: vibecoding
【HN】热度: 120 分 | 63 评论
【HN】热度: 120 分 | 63 评论