X / Twitter
Sam Altman (sama on X) — CEO, OpenAI OpenAI CEO Sam Altman said GPT-5.6 Sol's growth "is insane" and credited the inference team with "heroic work" to support demand — while warning that even as OpenAI moves mountains to keep scaling, "it is possible there are some hiccups soon." In a separate one-liner the same day, he called that situation "also, a reason to favor open-source harnesses."
OpenAI CEO Sam Altman 说 GPT-5.6 Sol 的增长"疯狂到不可思议",并称赞推理团队为支撑需求付出了"英雄般的努力"。他同时预警:即便 OpenAI 会不惜一切代价继续扩容,"近期仍有可能出现一些小故障"。同一天他还发了一句耐人寻味的短推:"这也是更倾向开源 harness 的一个理由。"
Thibault Sottiaux (thsottiaux on X) — Codex & ChatGPT, OpenAI OpenAI's Codex and ChatGPT lead Thibault Sottiaux teased that usage looks set to hit 9M soon ("Embarrassment of riches") and polled users on whether to reset ChatGPT Work and Codex usage limits again or hold off. He also spent the evening scrolling X for product feedback, asking directly: "What should we improve" in ChatGPT Work?
OpenAI Codex 与 ChatGPT 负责人 Thibault Sottiaux 预告用量看起来很快要突破 9M(他形容这是"幸福的烦恼"),并向用户发起投票:要不要再次重置 ChatGPT Work 和 Codex 的用量额度,还是先缓一缓。他当晚还在 X 上刷帖收集产品反馈,直接开问:ChatGPT Work"我们应该改进什么?"
Aaron Levie (levie on X) — CEO, Box Box CEO Aaron Levie argued that code is uniquely agent-friendly because you can test it quickly, whereas most other work only gets "tested" when it hits the real world: a trade executes, a contract gets negotiated, a pitch gets delivered. He predicts a whole new opportunity space around bringing that testability to the rest of work, noting most enterprise workflows today have no evals at all, and that "the enterprises that are able to eval their knowledge work the best also stand to gain the most from AI." He also backed a proposal for an AI standards body as distinct from a regulatory agency, warning that if AI progress moves at classic government speed, "America likely just loses the AI race."
Box CEO Aaron Levie 提出,代码之所以特别适合 agent,是因为它可以被快速验证;而大多数其他工作只有在真正进入现实世界时才能得到"检验":一笔交易被执行、一份合同谈成、一场销售演示讲完。他预判,把这种可测试性带到其他类型工作中,会催生一整片新的机会空间。他指出,如今企业里的绝大多数工作流根本没有任何 eval,而"最擅长对自己知识型工作做 eval 的企业,也将从 AI 中获益最多"。他还表态支持一项设立 AI 标准机构(而非监管机构)的提案,并警告:如果 AI 进步的速度退化到政府部门的常规节奏,"美国大概率会直接输掉这场 AI 竞赛"。
Guillermo Rauch (rauchg on X) — CEO, Vercel Vercel CEO Guillermo Rauch announced that Vercel is opening up its dataset of AI token flows on the Vercel AI Gateway — "Fascinating insights contained within!" — and shared a project he built on top of it. He also spotlighted AgentMail: telling your agent to run "vercel install agentmail" gets it email with no signup, automatic setup, and unified billing.
Vercel CEO Guillermo Rauch 宣布 Vercel 将开放其 AI Gateway 上的 AI token 流量数据集,"里面藏着非常有意思的洞察!",他还展示了自己基于这个数据集做的项目。他同时推荐了 AgentMail:让你的 agent 执行 "vercel install agentmail",就能获得无需注册、自动配置、统一计费的 agent 专用邮箱。
Ryo Lu (ryolu_ on X) — Design, Cursor Cursor design lead Ryo Lu published a reflective essay, "when the dream becomes the job," about what happens when the machines get good at the parts of creative work that "used to feel close to the heart" — writing, coding, designing, taste-like decisions. His answer to that fear: AI "cannot want on your behalf. it cannot decide what is worth loving. it cannot protect the strange private thread that made you care in the first place." The real discipline, he writes, is keeping a relationship with "the source" — the private part of the craft with no audience, roadmap, or deadline — because losing it leaves work that is "thinner. safer. more explainable. less alive."
Cursor 设计负责人 Ryo Lu 发表了一篇长文《当梦想变成工作》,探讨当机器开始擅长创造性工作中"曾经离内心最近"的那些部分(写作、编程、设计、带品味的决策)时会发生什么。他对这种恐惧的回答是:AI "无法替你产生渴望,无法替你决定什么值得热爱,也无法守护当初让你在乎这件事的那根隐秘的线"。他写道,真正的修行是与"源头"保持联结,也就是手艺中那个没有观众、没有 roadmap、没有 deadline 的私人角落;一旦失去它,作品就会变得"更单薄、更安全、更好解释,也更没有生命力"。
Thariq (trq212 on X) — Claude Code, Anthropic Anthropic's Thariq (Claude Code team) has been using Claude Code as a Pokemon Champions coach: it writes code against Smogon's npm library, pulls live usage stats, and generates reports to understand matchups, breakpoints, and theorycraft teams. He published his first public artifact — a breakdown of his Mega Sceptile team — and says he'll open-source the setup if there's interest.
Anthropic Claude Code 团队的 Thariq 最近把 Claude Code 当成了 Pokemon Champions 的教练:它调用 Smogon 的 npm 库写代码、拉取实时使用率数据,再生成报告来分析对局、伤害临界点和队伍构筑理论。他发布了自己的第一个公开 artifact(一份 Mega Sceptile 队伍拆解),并表示如果大家感兴趣就把整套方案开源。
Aditya Agarwal (adityaag on X) — GP, South Park Commons South Park Commons GP (and former Dropbox CTO) Aditya Agarwal flagged a real tradeoff in the new ChatGPT app: he appreciates the in-depth feature set, but he used the legacy app 15-20 times a day for quick queries, and the new version "feels so heavyweight for that use case. Shame."
South Park Commons 普通合伙人(前 Dropbox CTO)Aditya Agarwal 指出了新版 ChatGPT app 的一个真实取舍:他认可新版功能的深度,但他过去每天要用旧版 app 做 15 到 20 次快速查询,而新版"对这种使用场景来说实在太重了。可惜"。
Claude (claudeai on X) — Anthropic Anthropic's Claude account detailed Claude for Teachers, built for K-12 privacy: conversations are never used for model training, and student information is protected by a data processing agreement written to comply with FERPA. Ask for a lesson plan and Claude starts from your state standards and high-quality curricula by connecting through Learning Commons, then drafts a plan and student-facing materials teachers can revise and take into class.
Anthropic 的 Claude 官方账号详细介绍了 Claude for Teachers,一款为 K-12 隐私要求而设计的产品:对话永远不会被用于模型训练,学生信息受一份按 FERPA 合规标准撰写的数据处理协议保护。向它要一份教案,Claude 会通过 Learning Commons 接入你所在州的课程标准和高质量教材,然后起草教案和面向学生的材料,教师可以修改后直接带进课堂。
Podcasts
Training Data — Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden
The Takeaway: The next frontier of AI platforms isn't a smarter harness — it's "strategies," a coordination layer where tokens are treated as non-fungible workers that get different jobs: advising, executing, reflecting, grading.
Katelyn Lesse and Angela Jiang run Anthropic's platform — both the external developer APIs and the internal infrastructure Anthropic's own products sit on, deliberately kept as one set of primitives rather than bifurcated. They describe the platform as a three-layer cake: knowledge (the Messages API, tools, skills, memory), execution (a low-level harness plus managed infrastructure — today's Claude Managed Agents), and coordination, which is where the roadmap is headed.
The most contrarian thread: steering scaffolds are dying. Two years ago a harness was walls built to keep a model moving from point A to point B; models are now steerable enough that they encourage customers to delete that part entirely. What stays valuable is the verification logic between model and execution — the make-or-break detail in high-consequence domains like legal and finance — and the meta level above it: "You could take that same token and choose to actually reflect on your past agentic sessions and write learnings to memory so that the next agent does a good job."
The alpha is in strategies that everyone can describe but few can ship. Best-of-N, for example: the papers exist, but "to actually build that thing and put it into production so you can actually test it on users... that's really, really freaking hard." Anthropic gives itself maybe five token "jobs" internally and expects the ecosystem to invent 100,000 more. Notably, they aren't precious about whose infrastructure any of this runs on: self-hosted sandboxes ship via partnerships with Modal, Vercel, Cloudflare, and Amazon's new micro VMs, plus MCP tunnels that punch through corporate firewalls. And on token-cost anxiety, the advice is to rationalize rather than clamp down — route tasks by complexity within a model family, because capping AI usage outright "is kind of the wrong move." The ecosystem bet is summed up by their electricity analogy: the technology transforms the world only because everyone can wire into it, "and that's not something that anybody can do by themselves."
一句话要点:AI 平台的下一个前沿不是更聪明的 harness,而是"策略"(strategies):一个把 token 当作不可互换的工人来分派不同工种的协调层,有的负责建议,有的负责执行,有的负责反思,有的负责评分。
Katelyn Lesse 和 Angela Jiang 共同执掌 Anthropic 的平台:既包括对外的开发者 API,也包括 Anthropic 自家产品所依赖的内部基础设施。两者被刻意保持为同一套原语,而不是拆成两条线。她们把平台描述成一个三层蛋糕:知识层(Messages API、工具、skills、记忆)、执行层(低层 harness 加托管基础设施,也就是今天的 Claude Managed Agents),以及协调层,那正是路线图的去向。
最反直觉的一条主线是:转向式脚手架正在消亡。两年前的 harness 是为了让模型从 A 点走到 B 点而砌的两堵墙;如今模型的可控性已经好到她们会主动建议客户把这部分直接删掉。真正保值的,是模型与执行之间的验证逻辑,这在法律、金融这类高后果领域是成败攸关的细节;再往上则是元层面:"你可以拿同一个 token,选择让它回顾过去的 agent 会话、把经验写进记忆,让下一个 agent 干得更好。"
超额收益藏在人人都会讲、却很少有人能落地的策略里。以 best-of-N 为例:论文早就有了,但"要真正把它构建出来、放进生产环境、在真实用户身上验证效果……那是真的非常非常难"。Anthropic 内部给 token 定义的"工种"也就五个左右,却预期生态能发明出十万个。值得注意的是,她们对这一切跑在谁的基础设施上并不执着:自托管沙箱通过与 Modal、Vercel、Cloudflare 以及 Amazon 新的 micro VM 的合作交付,还有能穿透企业防火墙的 MCP 隧道。至于 token 成本焦虑,她们的建议是理性化而不是一刀切:在同一个模型家族内按任务复杂度做路由,因为直接叫停 AI 使用"基本上是个错误的动作"。整个生态的赌注可以用她们的电力类比来概括:这项技术之所以能改变世界,正是因为所有人都能接入电网,"而这不是任何一家公司能独自做到的事"。