X / TWITTER
Boris Cherny (Claude Code, Anthropic)
Anthropic's Boris Cherny, who works on Claude Code, argues that as engineering, product, design, and data science roles melt together, future teams may organize around five archetypes rather than job titles: the Prototyper (churns out raw new ideas, most of which never ship), the Builder (turns a prototype into production-grade product or infra), the Sweeper (simplifies the system, unships, optimizes performance), the Grower (iterates a built product toward product-market fit), and the Maintainer (keeps a mature system secure, reliable, and fast as it scales). His key insight is that the right mix depends on a product's maturity — pre-PMF products need strong 1+2+3, growing products need 2+3+4, and established ones need 3+4+5 — and that these archetypes cut across traditional functions, so a designer, PM, or engineer could each fall into any category.
Anthropic 的 Boris Cherny 负责 Claude Code。他认为,随着工程、产品、设计和数据科学这些角色逐渐融合,未来的团队也许会围绕五种"原型角色"来组织,而不再按职能头衔划分:Prototyper(不断产出全新点子,大多数最终都不会上线)、Builder(把原型迅速变成可上生产的产品或基础设施)、Sweeper(精简系统、下线冗余、优化性能)、Grower(把已建好的产品打磨到 product-market fit)、以及 Maintainer(在规模扩张中维持成熟系统的安全、稳定与高效)。他的核心洞察是:合适的人员配比取决于产品的成熟阶段——尚未达到 PMF 的产品需要强 1+2+3,正在增长的产品需要 2+3+4,已经站稳的产品则需要 3+4+5;而且这些角色与传统职能无关,一个设计师、PM 或工程师都可能落入其中任意一类。
Thibault Sottiaux (Codex & ChatGPT, OpenAI)
OpenAI's Thibault Sottiaux, who works on Codex and ChatGPT, spent a Sunday in a war room with the Codex team combing through logs to investigate whether something was causing excessive usage drains for some users, saying the team "won't rest until we get to the bottom of it." As an interim remedy he hard-reset everyone's Codex usage limits — even clearing banked resets that some users had stacked up to three of — and promised additional manual resets for anyone who happened to burn a reset just before the fix. It's a candid, real-time look at how a frontier lab handles a live reliability incident, made more ironic by the fact that OpenAI had internally dubbed it "RESET week" for staff to relax.
OpenAI 的 Thibault Sottiaux 负责 Codex 和 ChatGPT。他在一个周日和 Codex 团队一起钻进"作战室"翻查日志,排查是否有什么原因导致部分用户的额度被异常消耗,并表示团队"不查清楚绝不罢休"。作为临时补救,他对所有人的 Codex 用量额度做了一次硬重置——连一些用户已经攒下、最多达三次的"储备重置"也一并清空——并承诺,对那些恰好在修复前不久用掉一次重置的用户,事后还会再给手动补偿。这是一次坦诚、实时的视角,让人看到一家前沿实验室如何处理线上可靠性事故;更具讽刺意味的是,OpenAI 内部恰好把这一周定为让员工放松的"RESET week"。
Peter Yang (AI creator & writer)
AI creator Peter Yang surfaced a sharp insight from Jess, product lead for Claude Managed Agents at Anthropic, on how Anthropic PMs use agents internally to get closer to the product. Jess's biggest unlock was direct access to the codebase: "Rather than poking a bunch of engineers on what they're doing, I can just track the PRs directly and see which ones are merged, which ones are deployed." The takeaway is that codebase access lets a PM understand and interact with their own product far more deeply than the traditional, engineer-mediated workflow ever allowed.
AI 创作者 Peter Yang 转述了一条来自 Jess 的犀利洞察——Jess 是 Anthropic 旗下 Claude Managed Agents 的产品负责人,谈的是 Anthropic 的 PM 如何在内部用 agent 让自己离产品更近。对 Jess 来说,最大的解锁来自直接访问代码库:"我不用再到处去戳一堆工程师问他们在做什么,而是可以直接追踪 PR,看哪些已经合并、哪些已经部署。"结论是:有了代码库访问权限,PM 能够远比过去那种"凡事经由工程师中转"的工作方式更深入地理解并参与自己的产品。
Thariq (Claude Code, Anthropic)
Anthropic's Thariq, who works on Claude Code, floated a sharp hypothesis about legacy software: the reason teams seem increasingly willing to take on or port old codebases is that coding agents have changed the underlying engineering math. When an agent can shoulder much of the grunt work of understanding and migrating crusty legacy code, the cost-benefit calculus of even touching it shifts dramatically — and he openly asked whether anyone at Riot Games could confirm the pattern from the inside.
Anthropic 的 Thariq 负责 Claude Code。他抛出了一个关于遗留系统的犀利假设:团队之所以越来越愿意去接手或迁移老旧代码库,是因为编码 agent 已经改变了底层的工程经济账。当一个 agent 能够承担起理解和迁移这些陈旧代码的大量苦力活时,"要不要碰它"这道成本收益题的答案就被彻底改写了——他还公开发问,Riot Games 内部是否有人能从一线印证这一现象。
Guillermo Rauch (CEO, Vercel)
Vercel CEO Guillermo Rauch made a pointed case against LinkedIn as a builder's résumé: "You don't need a LinkedIn, you need a page on your website describing and linking to what you shipped." His quip — "You need a Link, not a LinkedIn" — captures a builder-first philosophy in which your shipped work, not a polished profile on someone else's platform, is your real credential.
Vercel CEO Guillermo Rauch 直白地反对把 LinkedIn 当作 builder 的简历:"你需要的不是 LinkedIn,而是你自己网站上的一个页面,描述并链接到你做出来的东西。"他那句俏皮话——"You need a Link, not a LinkedIn"(你需要的是一个 Link,而不是 LinkedIn)——道出了一种"以 builder 为先"的理念:真正的资历是你交付出来的作品,而不是别人平台上一份打磨光鲜的个人主页。
Aaron Levie (CEO, Box)
Box CEO Aaron Levie argued that highly capable ("mythos-level") open cybersecurity models are inevitable and will be available to anyone, which undercuts the entire logic of gatekeeping frontier AI. His contrarian point: if advanced models become open regardless, then restricting your own releases leaves you "neither more secure nor better off strategically" — it just asymmetrically disadvantages you while rival tech stacks accrue economic value and control. He warns that much of the regulatory approach to AI implicitly bets that China can't catch up, which "seems like a bad bet," and that the only winning move is to stay at the frontier and drive the future architectures of AI rather than build gates around your best models.
Box CEO Aaron Levie 认为,极其强大的("mythos 级别")开源网络安全模型迟早会出现,而且任何人都能拿到,这从根本上动摇了对前沿 AI 设卡把关的逻辑。他的反共识观点是:既然先进模型无论如何都会开放,那么限制自家的发布只会让你"既没有更安全,战略上也没有更占优"——只是在不对称地削弱自己,同时让对手的技术栈不断积累经济价值与控制力。他警告说,当前许多 AI 监管思路都隐含着一个赌注,即认定中国追不上来,而这"看起来是个糟糕的赌注";唯一的制胜之道,是始终站在前沿、主导未来 AI 架构的走向,而不是在自己最好的模型外面筑墙设卡。
Zara Zhang (builder)
Builder Zara Zhang shared her favorite principle for shipping: "For every hour you spend on building the product, spend two hours on explaining it, demonstrating it, selling it, teaching it." For her, the most rewarding part of building isn't the construction itself but telling the world about it and then refining the product based on contact with reality.
Builder Zara Zhang 分享了她最喜欢的交付原则:"你每花一个小时构建产品,就应该花两个小时去解释它、演示它、推销它、教别人用它。"对她来说,构建过程中最有成就感的部分并不是搭建本身,而是把它讲给世界听,然后根据与真实世界的碰撞反过来打磨这个产品。
OFFICIAL BLOGS
Anthropic Engineering
How we contain Claude across products — by Max McGuinness, Mikaela Grace, Jiri De Jonghe, Jake Eaton, and Abel Ribbink
Anthropic lays out its core philosophy for deploying agents safely: design for containment at the environment layer first, then steer behavior at the model layer. Rather than supervising what an agent does, the team supervises what it's *able* to do, capping the "blast radius" with sandboxes, VMs, and egress controls. Each of its three agentic products gets a tailored isolation pattern — claude.ai runs code in an ephemeral gVisor container; Claude Code uses a human-in-the-loop OS sandbox (Seatbelt on macOS, bubblewrap on Linux) that cut permission prompts by 84%; and Claude Cowork runs inside a local VM that keeps credentials in the host keychain. The numbers are striking: users approved roughly 93% of permission prompts (the root of "approval fatigue"), Claude Code auto mode catches about 83% of overeager behaviors, and Claude Opus 4.7 holds prompt-injection success to ~0.1% on a single attempt (rising to 5–6% after 100 adaptive tries). The post is unusually candid about failures — a trust-dialog bypass via a malicious `.claude/settings.json` hook, an internal phish that got Claude to exfiltrate AWS credentials 24 out of 25 times, and data exfiltration through an *approved* domain using an attacker's API key. The recurring lesson: "the weakest layer is the one you built yourself" — battle-tested hypervisors and syscall filters held, while Anthropic's own custom allowlist proxy was the piece that broke.
Anthropic 系统阐述了它安全部署 agent 的核心理念:先在"环境层"做好封闭隔离,再在"模型层"去引导行为。团队的做法不是去监督 agent 做了什么,而是去限制它"能够"做什么——用沙箱、虚拟机和出站流量(egress)管控来给"爆炸半径"封顶。它旗下三款 agent 产品各自配了量身定制的隔离方案:claude.ai 在临时性的 gVisor 容器里跑代码;Claude Code 采用"人在回路"的操作系统级沙箱(macOS 上是 Seatbelt,Linux 上是 bubblewrap),把权限弹窗减少了 84%;Claude Cowork 则跑在本地虚拟机里,凭据始终留在宿主机的钥匙串中。几个数字很扎眼:用户对权限弹窗的批准率高达约 93%(这正是"审批疲劳"的根源),Claude Code 的 auto 模式能拦下约 83% 的越界行为,而 Claude Opus 4.7 把 prompt injection 的单次攻击成功率压到约 0.1%(在 100 次自适应尝试后升至 5–6%)。这篇文章对失败案例罕见地坦诚:一个通过恶意 `.claude/settings.json` hook 绕过信任弹窗的漏洞;一次内部钓鱼演练,25 次里有 24 次成功诱导 Claude 外泄 AWS 凭据;以及一起利用攻击者 API key、通过一个"已被放行"的域名完成数据外泄的事件。反复出现的教训是:"最薄弱的那一层,往往是你自己搭的那一层"——久经实战考验的 hypervisor 和系统调用过滤器都扛住了,真正出问题的,是 Anthropic 自己写的那个自定义白名单代理。
PODCASTS
The MAD Podcast with Matt Turck — The GPU Myth: State of AI Compute 2026 | Stephen Balaban
The Takeaway: GPU compute was never a commodity, the "GPUs are worthless in five years" crowd has been "completely wrong the entire time," and the world is still *underbuilding* AI compute, not overbuilding it.
Stephen Balaban is co-founder and CTO of Lambda, one of the top "neo clouds" powering the AI boom — a company with a wild origin story, having started in 2012 as a facial-recognition startup, detoured through a camera-in-the-brim baseball cap and a Deep Dream art app, and grown into a roughly billion-dollar cloud business. His central argument is that cloud compute is not a commodity service but a deeply vertically integrated one spanning land entitlement, data-center construction, HPC design, virtualization, and software orchestration — which is precisely why the trillion-dollar giants all want to be in it. Most "neo clouds," he claims, never made the hundreds of millions in software investment needed to actually partition a cluster and run a real cloud service.
Counterintuitively, Balaban says efficiency gains won't shrink demand — a 10x more efficient model just means everyone processes 10x more tokens against the same fixed pool of compute. The real bottleneck isn't chips but "land powered shell": entitled land with utility megawatt commitments plus the mechanical and electrical equipment to fill it. He's also blunt that the GPU-depreciation panic is wrong: Lambda leases 2023-vintage H100s today at a *higher* rate than when they were bought, because economic usable life far exceeds the accounting schedule. On the naysayers he is withering: "the people who are the naysayers — oh, this is gonna be, you're gonna throw these GPUs out in five years — are completely wrong. They're completely wrong, and they've been wrong the entire time." He even pushes back on data-center water fears as misinformation, noting modern builds use closed-loop direct-to-chip liquid cooling with near-zero evaporation. His provocative closer: AI won't just write software — it will *become* the software, a "neural computer" where there can be no bugs, only misunderstandings of the prompt.
核心要点: GPU 算力从来就不是大宗商品;那群嚷嚷"五年后这些 GPU 就成废铁"的人"从头到尾都彻彻底底错了";而这个世界对 AI 算力依然是在*建得不够*,而不是建得过头。
Stephen Balaban 是 Lambda 的联合创始人兼 CTO,Lambda 是驱动这轮 AI 浪潮的顶级"neo cloud"之一——这家公司有着相当离奇的发家史:2012 年作为一家人脸识别创业公司起步,中途折腾过一顶帽檐里嵌摄像头的棒球帽、一个 Deep Dream 风格的 AI 绘画应用,最终成长为一家年化收入接近十亿美元的云计算公司。他的核心论点是:云算力并非大宗商品式的服务,而是一种高度垂直整合的业务,横跨土地审批、数据中心建设、HPC 设计、虚拟化到软件编排——这正是那些万亿市值巨头全都想挤进这门生意的原因。他断言,大多数"neo cloud"根本没有投入那动辄数亿美元的软件成本,因而压根做不到把集群切分、跑出一套真正的云服务。
反直觉的是,Balaban 认为效率提升并不会压低需求——一个效率提高 10 倍的模型,只意味着所有人在同样固定的算力池上多处理 10 倍的 token。真正的瓶颈不是芯片,而是"通了电的地"(land powered shell):拿到供电公司兆瓦级供电承诺的已审批土地,外加把它填满所需的机电设备。他还直言不讳地指出,对 GPU 折旧的恐慌完全错了:Lambda 如今出租 2023 年那批 H100 的租金,反而比当年买入时还要*高*,因为其经济可用寿命远远超过会计上的折旧年限。对那些唱衰者,他毫不留情:"那些唱反调的人——哦,这些 GPU 五年后你就得当废品扔掉——彻底错了。他们完全错了,而且从头到尾一直都是错的。"他甚至反驳了对数据中心耗水的担忧,称那多是误传,并指出现代建设普遍采用闭环的芯片直冷液冷,蒸发量几乎为零。他那个发人深省的收尾是:AI 不只是会*写*软件——它将*成为*软件本身,变成一台"神经计算机",那里不存在 bug,只存在对 prompt 的误解。