X / Twitter
Swyx (Latent Space / Cognition)
Swyx of Cognition and Latent Space says the most underpriced growth hack right now is pointing your coding agents at your own marketing: if you haven't set your Codex, Claude, Gemini, or Devin automations to auto-research how to improve your SEO/AEO every week, "you are really truly missing out on free, should-be-commoditizing-but-weirdly-untapped alpha." He adds that the alpha only disappears once the conversation matures past basic tips into questions like on-policy AEO — does Claude optimizing your AEO disproportionately improve how Claude itself ranks you — versus generalizable AEO.
Cognition 和 Latent Space 的 Swyx 认为,眼下最被低估的增长手段就是让你的 coding agent 盯着自家的营销做优化:如果你还没让 Codex、Claude、Gemini 或 Devin 的自动化任务每周自动研究怎么改进你的 SEO/AEO,"你真的在白白错过一份免费的、本该早就被用烂却奇怪地没人挖的 alpha"。他补充说,这份 alpha 只有等讨论从基础技巧升级到 on-policy AEO 这类问题时才会消失,比如:让 Claude 优化你的 AEO,会不会不成比例地提升 Claude 自己给你的排名,还是说存在可泛化的 AEO。
Thibault Sottiaux (OpenAI, Codex & ChatGPT)
OpenAI's Codex and ChatGPT lead Thibault Sottiaux reset usage limits for all paid Codex and ChatGPT Work users — "Oops... I did it again" — crediting a team "iterating at lightspeed and keeping the infra up as we scale faster than ever," and cheekily noting the move "might also have reset other rate limits out there. Let's see. You're welcome if so." He also quote-tweeted Anthropic's Fable capacity announcement with the quip "GPT-5.6 Sol confirmed to be an extremely good model."
OpenAI 负责 Codex 和 ChatGPT 的 Thibault Sottiaux 宣布为所有 Codex 和 ChatGPT Work 付费用户重置了用量额度,配文"Oops... I did it again",并感谢团队"以光速迭代,在我们以前所未有的速度扩张时稳住了 infra",还俏皮地表示这波操作"可能连带重置了别家的限额,走着瞧,如果是的话不用谢"。他还转评了 Anthropic 关于 Fable 容量的公告,调侃道"GPT-5.6 Sol 被证实是个非常好的模型"。
Claude (Anthropic official)
Anthropic's Claude account announced that beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans at 50% of limits, while Pro and Team Standard users keep access via usage credits plus a one-time $100 credit. Anthropic acknowledged that "demand for Fable has been challenging to predict," which is why it rolled out to subscription plans in stages, and says it is continuing to invest in new capacity.
Anthropic 的 Claude 官方账号宣布,从 7 月 20 日起,Claude Fable 5 将纳入所有 Max 和 Team Premium 套餐,按 50% 限额提供;Pro 和 Team Standard 用户继续通过 usage credits 使用 Fable,并将获得一次性 100 美元额度。Anthropic 承认"Fable 的需求一直很难预测",这也是它分阶段向订阅套餐开放的原因,并表示会继续投入新的容量建设。
Madhu Guru (Meta, Sr Director of AI)
Meta's Senior Director of AI Madhu Guru pushed back on the idea that Kimi hurts Google: many enterprises won't consume Kimi directly, they'll get it through Google Cloud because they still need enterprise guarantees — security, data residency, compliance, and most importantly chips. "Money out one pocket into the other." He also argued the reason enterprises struggle to go beyond basic chatbots is a talent gap in evals and harnesses: evals that express your ambition and push the jagged frontier of the models, a harness that manages routing, multi-agent orchestration, context management, tool calling and memory independently of the models — and, scarcest of all, the talent to build both at the frontier.
Meta AI 高级总监 Madhu Guru 反驳了"Kimi 会伤到 Google"的说法:很多企业不会直接用 Kimi,而是会通过 Google Cloud 来用,因为它们仍然需要企业级保障,安全、数据驻留、合规,以及最重要的芯片。"钱从一个口袋出来,进了另一个口袋。"他还指出,企业迈不过基础 chatbot 这道坎,根源是构建 evals 和 harness 的人才缺口:evals 要能表达你的野心、推动模型的锯齿状能力边界;harness 要独立于模型,管好路由、多 agent 编排、context 管理、tool calling 和 memory;而最稀缺的,是能在前沿把这一切搭起来的人才。
Aaron Levie (Box CEO)
Box CEO Aaron Levie argued that cheaper AI benefits the entire ecosystem, especially end customers: everything is bottlenecked by deploying AI cost-effectively in real workloads, so every price drop pushes total usage up and value accrues to all layers of the stack. His twist: efficiency gains can ironically increase spend on frontier closed models too, because you often need the strongest model to orchestrate a task while farming out the bulk of tokens to cheaper or tuned models. The thing most at risk is margins, which he expects to eventually converge with infrastructure-stack margins.
Box CEO Aaron Levie 认为 AI 变便宜会让整个生态受益,尤其是终端客户:一切都卡在能否把 AI 低成本地部署进真实工作负载上,所以每次降价都会推高总用量,价值会沉淀到技术栈的每一层。他的反直觉观点是:效率提升反而可能让前沿闭源模型的开销更大,因为你往往需要最强的模型来编排任务,再把大头 token 分发给更便宜或更专门调优的模型。最危险的是利润率,他预计智能的利润率最终会向基础设施层的利润率收敛。
Peter Steinberger (OpenClaw)
OpenClaw's Peter Steinberger described watching Codex use browser and computer use to open Chrome, navigate to his GitHub PR, and wrangle the macOS file picker just to upload an image — because GitHub has no API for it — calling it "both amazing and painful"; he now runs his Codex agents in VMs so they don't steal app focus. After endless user requests around CodexBar icon customization, he also had Codex build a full icon editor. And he captured the agent-architecture zeitgeist in one line: "Are we still talking loops or did we shift to graphs yet?"
OpenClaw 的 Peter Steinberger 描述了他看着 Codex 用浏览器加 computer use 打开 Chrome、进到他的 GitHub PR、再跟 macOS 文件选择器搏斗,只为上传一张图片,原因是 GitHub 没有对应的 API。他形容这个过程"既惊艳又痛苦",现在他把 Codex agent 都跑在虚拟机里,免得它们抢走应用焦点。被 CodexBar 图标定制的需求轰炸之后,他还让 Codex 做了一个完整的图标编辑器。他还用一句话概括了当下 agent 架构的讨论氛围:"我们还在聊 loop 吗,还是已经转向 graph 了?"
Thariq (Anthropic, Claude Code)
Anthropic Claude Code team member Thariq shared a token-saving workflow tip: building prototypes of mockups, schemas, data models, and proofs of concept first is the best way to avoid spending tons of tokens before realizing you don't actually want the output.
Anthropic Claude Code 团队的 Thariq 分享了一个省 token 的工作流技巧:先做 mockup、schema、数据模型和 proof of concept 的原型,是避免烧掉大量 token 之后才发现产出根本不是你想要的东西的最好办法。
Peter Yang (Creator Economy)
Creator Economy author Peter Yang wants agent management to go hands-free: spending an entire day staring at screens managing agents "fries my brain," and he'd much rather be walking around outside talking to agents "on the phone," assigning them work and getting status updates by voice. "Can't wait for the first lab to ship this."
Creator Economy 作者 Peter Yang 希望 agent 管理能彻底解放双手:一整天盯着屏幕管 agent"会把我的脑子烤糊",他更想在外面边散步边跟 agent"打电话",口头派活,让它们用语音汇报进度。"等不及哪家实验室先把这个做出来了。"
Zara Zhang (Builder)
Builder Zara Zhang shared a tip for building in public: if making content feels like extra work, show the work already happening inside your product — a tiny screen recording, the first version, or the user behavior that changed your design — because the reasoning matters more than production value. She also observed how fast technology changes culture: just a few years ago many people were uncomfortable with recording meetings, and now it's simply assumed all business meetings are recorded, not for humans but for agents.
Builder Zara Zhang 分享了一个 build in public 的技巧:如果做内容让你觉得是额外负担,那就直接展示产品里正在发生的工作,比如一段小小的录屏、第一个版本、或者促使你改设计的用户行为,因为思考过程比制作精良更重要。她还感慨技术改变文化之快:几年前很多人还对会议录音感到不适,如今大家默认所有商务会议都会被录下来,不是给人看的,是给 agent 用的。
Guillermo Rauch (Vercel CEO)
Vercel CEO Guillermo Rauch announced that Sandbox data for downloads is now free, telling builders "Happy Friday! Time to ship more agents."
Vercel CEO Guillermo Rauch 宣布 Sandbox 的下载流量现在免费了,并对开发者说"周五快乐!该去 ship 更多 agent 了"。
Official Blogs
Anthropic Engineering: An update on recent Claude Code quality reports
Anthropic published a candid postmortem tracing weeks of "Claude got worse" reports to three separate shipped changes affecting Claude Code, the Claude Agent SDK, and Claude Cowork — the API was never impacted, and all three were resolved as of April 20 (v2.1.116). First, a March 4 change dropped Claude Code's default reasoning effort from high to medium to cut latency; users felt the intelligence loss and Anthropic reverted it, now defaulting to xhigh for Opus 4.7 and high elsewhere. Second, a March 26 caching optimization meant to clear old thinking once from stale sessions instead cleared it every turn due to a bug, making Claude forgetful and repetitive while draining usage limits through cache misses; it survived multiple reviews, tests, and dogfooding, and took over a week to root-cause. Notably, in back-testing, Opus 4.7's code review caught the bug while Opus 4.6 didn't. Third, an April 16 system prompt line capping response lengths to reduce verbosity caused a 3% drop on one eval and was reverted. "We take reports about degradation very seriously. We never intentionally degrade our models." Going forward: more staff on exact public builds, per-line prompt ablations with per-model evals, soak periods and gradual rollouts — plus a usage limit reset for all subscribers.
Anthropic 发布了一篇坦诚的事后复盘,把持续数周的"Claude 变笨了"反馈追溯到三个分别上线的改动,波及 Claude Code、Claude Agent SDK 和 Claude Cowork,API 始终未受影响,三个问题已在 4 月 20 日(v2.1.116)全部修复。第一,3 月 4 日为降低延迟把 Claude Code 默认推理强度从 high 调到 medium,用户明显感到智能下降,Anthropic 随后回滚,现在 Opus 4.7 默认 xhigh,其他模型默认 high。第二,3 月 26 日一项缓存优化本想只对闲置会话清理一次旧的 thinking,却因为 bug 变成每一轮都清,导致 Claude 显得健忘、重复,还因 cache miss 加速消耗用量额度;这个 bug 通过了多轮人工和自动审查、测试和内部试用,花了一周多才定位。值得一提的是,回测中 Opus 4.7 的 code review 抓到了这个 bug,而 Opus 4.6 没有。第三,4 月 16 日 system prompt 里一条限制回复长度的降冗余指令在某项评测上造成 3% 下降,已回滚。"我们非常严肃地对待关于降智的反馈。我们从不会故意让模型变差。"后续措施包括:让更多员工使用与公众完全一致的版本、对 prompt 逐行做分模型评测消融、增加浸泡期和灰度发布,同时为所有订阅用户重置用量额度。
Anthropic Engineering: Scaling Managed Agents: Decoupling the brain from the hands
Anthropic explains the architecture behind Managed Agents, its hosted service for long-horizon agents, built on an old computing insight: design interfaces for "programs as yet unthought of," the way operating systems virtualized hardware into processes and files. The team virtualized the agent into a session (append-only event log), a harness (the loop calling Claude), and a sandbox (execution environment), each swappable independently. The original everything-in-one-container design had become a "pet" — a hand-tended server whose failure lost the session — so they decoupled the brain (Claude and harness) from the hands (sandboxes and tools): containers became disposable cattle provisioned on demand via tool calls, credentials moved out of the sandbox entirely (bundled with resources or held in a vault behind an MCP proxy), and the durable session log lets a crashed harness reboot and resume. The payoff was concrete: p50 time-to-first-token dropped roughly 60% and p95 over 90%. A key theme is that harnesses encode assumptions that go stale as models improve — context resets built for Sonnet 4.5's "context anxiety" became "dead weight" on Opus 4.5 — so Managed Agents is deliberately a meta-harness, "opinionated about the shape of these interfaces, not about what runs behind them."
Anthropic 详解了 Managed Agents 这项长周期 agent 托管服务背后的架构,出发点是计算机领域的一个古老洞见:要为"尚未被想到的程序"设计接口,就像操作系统把硬件虚拟化成进程和文件那样。团队把 agent 虚拟化为三个部件:session(只追加的事件日志)、harness(调用 Claude 的循环)和 sandbox(执行环境),三者可以独立替换。最初的单容器设计变成了一只"宠物",一台需要人工照料、一挂就丢会话的服务器;于是他们把"大脑"(Claude 和 harness)与"双手"(sandbox 和工具)解耦:容器变成按需通过 tool call 创建、可随意丢弃的"牲口",凭证被彻底移出 sandbox(要么与资源捆绑,要么存在 MCP 代理背后的保险库里),持久化的 session 日志让崩溃的 harness 可以重启续跑。收益非常具体:首 token 时间 p50 下降约 60%,p95 下降超过 90%。文章的核心主题是 harness 会固化对模型能力的假设,而这些假设会随模型进步而过时,比如为 Sonnet 4.5 的"context 焦虑"设计的 context 重置,到了 Opus 4.5 上就成了"死重"。所以 Managed Agents 刻意做成一个元 harness,"我们只对接口的形状有主见,对接口背后跑什么没有主见"。
Claude Blog: New in Claude Managed Agents: self-hosted sandboxes and MCP tunnels
Claude Managed Agents can now execute inside a sandbox you control and reach MCP servers on your private network. Self-hosted sandboxes (public beta) keep code execution, sensitive files, packages, and data inside your enterprise perimeter — on your own infrastructure or via managed providers Cloudflare, Daytona, Modal, or Vercel — while the agent loop handling orchestration, context management, and error recovery stays on Anthropic's infrastructure. Early adopters include Amplitude (an internal Design Agent on Cloudflare), Clay (its GTM engineering agent Sculptor on Daytona), and finance AI platform Rogo (an analyst agent on Vercel Sandbox, whose firewall injects credentials at the network boundary so they never enter the sandbox). MCP tunnels (research preview) connect agents to internal databases, private APIs, and ticketing systems via a lightweight gateway that makes a single outbound connection — no inbound firewall rules, no public endpoints, traffic encrypted end to end.
Claude Managed Agents 现在可以在你自己掌控的 sandbox 里执行,并访问你私有网络内的 MCP 服务器。自托管 sandbox(公测中)让代码执行、敏感文件、依赖包和数据都留在企业边界之内,可以跑在自有基础设施上,也可以用 Cloudflare、Daytona、Modal 或 Vercel 这些托管服务商;负责编排、context 管理和错误恢复的 agent 循环仍留在 Anthropic 的基础设施上。早期用户包括 Amplitude(基于 Cloudflare 的内部 Design Agent)、Clay(跑在 Daytona 上的 GTM 工程 agent Sculptor)和金融 AI 平台 Rogo(基于 Vercel Sandbox 的分析师 agent,Vercel 的防火墙在网络边界注入凭证,凭证永远不会进入 sandbox)。MCP tunnels(研究预览)通过一个轻量网关把 agent 连到内部数据库、私有 API 和工单系统,网关只发起一条出站连接,不需要入站防火墙规则,没有公开端点,流量端到端加密。
Podcasts
The MAD Podcast with Matt Turck — "OpenAI's Compute Chief: We Can't Build Fast Enough | Sachin Katti"
The Takeaway: The binding constraint on AI is not money or conviction but the physical world — OpenAI's compute chief fears building too little, never too much.
Sachin Katti runs "industrial compute" at OpenAI — a Stanford professor, multi-time founder, and former Intel CTO now leading what he calls one of the largest things humanity has ever built, with OpenAI spending directionally around $50 billion on compute this year inside an industry spending roughly $700 billion. His mental model: data centers are "giant factories that are turning electrons into tokens." When asked about overbuild risk, he flips the question. "Anytime you have thought you have enough compute, we can slow down. Always negatively surprises like, oh, we should not have slowed down. Demand far outstrips compute supply today." The pattern so far: tripled compute has meant tripled revenue, and every new cluster is consumed immediately. The training-versus-inference distinction is dissolving — synthetic data generation, post-training, and test-time compute are all inference, making inference the majority workload even during training. And because AI can now run AI research experiments, compute demand is no longer bounded by the scarce supply of human researchers. The most striking claim involves Jalapeno, the custom inference chip OpenAI designed with Broadcom in nine months — the fastest he's seen in his career — optimizing tokens per watt, partly because AI accelerated the design: "the world of recursion is not that far where AI will design the systems it needs to train and run the next generation of AI." The real bottlenecks are mundane: gas turbines, transformers, electricians, and plumbers, industries hit by a demand shock after a decade of flat capacity. Nuclear "can't come soon enough." Notably, OpenAI is always the tenant, never the owner — partners like Oracle, Microsoft, and SoftBank finance and build while OpenAI commits to consuming the tokens, and its new guaranteed-capacity offering treats intelligence itself as a supply unit enterprises lock in like any critical input.
核心要点:AI 的硬约束不是钱也不是信念,而是物理世界。OpenAI 的算力负责人担心的从来不是建多了,而是建少了。
Sachin Katti 在 OpenAI 负责"工业级算力",他是 Stanford 教授、连续创业者、前 Intel CTO,如今在主导他口中人类有史以来最大的建设工程之一:OpenAI 今年在算力上的支出大方向在 500 亿美元左右,整个行业约 7000 亿美元。他的心智模型是:数据中心就是"把电子变成 token 的巨型工厂"。被问到过度建设的风险时,他把问题反了过来:"每次我们觉得算力够了、可以放缓的时候,结果总是负面惊喜:哎呀,我们不该慢下来的。今天的需求远远超过算力供给。"目前的规律是:算力翻三倍,收入就翻三倍,每上线一个新集群都会被立刻用光。训练和推理的界限正在消失,合成数据生成、post-training、test-time compute 本质上都是推理,即便在训练阶段推理也占了大头。而且因为 AI 现在自己就能跑 AI 研究实验,算力需求不再受限于稀缺的人类研究员数量。最令人印象深刻的是 Jalapeno,OpenAI 与 Broadcom 合作、九个月内完成设计到流片的自研推理芯片,这是他职业生涯见过最快的速度,优化目标是每瓦 token 数,而提速的原因之一正是 AI 参与了芯片设计:"递归的世界已经不远了,AI 将设计出用来训练和运行下一代 AI 的系统。"真正的瓶颈反而很接地气:燃气轮机、变压器、电工和管道工,这些行业十年没扩产,突然被需求冲击砸中。核电"来得越快越好"。值得注意的是,OpenAI 永远只做租户、不做业主:Oracle、Microsoft、SoftBank 这些伙伴负责融资和建设,OpenAI 承诺消费 token;而它新推出的保障容量服务,本质是把智能当作一种供应品,让企业像锁定任何关键物资一样锁定 token 供给。