回到卷首
每日集录ai builders

七月二十四日

二〇二六年 12 builders 27 posts 1 podcast 1 blog 约三十五分钟

Cerebras CEO Andrew Feldman tells The MAD Podcast that inference speed became AI's bottleneck the moment models got useful in mid-2025 ('there's no moat in inference, it takes you eight keystrokes to move from a GPU to us'), that generating each token means moving 100 HD movies' worth of weights, and that the HBM/CoWoS/3nm shortages don't touch his SRAM wafer-scale chip; Anthropic's Claude Blog ships artifacts in Claude Code, turning sessions into live shareable pages private to your org; and on X, OpenAI ships voice in the ChatGPT desktop app (Codex lead Thibault Sottiaux hails 'Jarvis / Samantha / TARS'), Anthropic's Claude voice mode gets Opus/Sonnet and mid-conversation tool access, Swyx dogfoods an agentic GitHub clone, Replit CEO Amjad Masad cuts Autoscale deployment costs 80%, and Meta AI Sr Director Madhu Guru asks how identity management survives when one employee can spawn hundreds of agents.

X / TWITTER

Swyx (Cognition / Latent Space)

Swyx has spent the past month dogfooding an agentic GitHub clone and says it's "gotten quite quite enjoyable to use," complete with built-in CI/CD thanks to Workers for Platforms. Three more ideas remain before it goes live, and he's inviting people to join Swyx Inc to hack on it and influence the roadmap.

Sources1

Swyx 过去一个月一直在内部试用自己做的 agentic GitHub 克隆版,他说 "已经用得相当相当顺手了",还借助 Workers for Platforms 内置了 CI/CD。上线前还有三个想法要实现,他也在邀请感兴趣的人加入 Swyx Inc 一起开发、影响产品路线图。

He also called out Poolside's "unusual degree of openness": beyond shipping a small model that beat Thinking Machines at coding and publishing well-regarded papers, they're among the rare few to expose their full eval dataset, beautifully published across 6 public benchmarks with 4 runs each and hundreds of turns per run, so you can verify for yourself whether they reward-hack. His verdict: "brilliant."

Sources1

他还点赞了 Poolside "不同寻常的开放程度":不仅发布了一个在编程上击败 Thinking Machines 的小模型、论文广受好评,还是极少数完整公开 eval 数据集的团队,覆盖 6 个公开 benchmark,每个跑 4 轮、每轮上百个 turn,公开得非常漂亮,你可以自己验证他们有没有 reward hacking。他的评价是:"brilliant。"

OpenAI's Thibault Sottiaux (Codex & ChatGPT)

OpenAI shipped voice into the ChatGPT desktop app, and Codex lead Thibault Sottiaux framed it as the sci-fi assistant finally arriving: "Jarvis / Samantha / TARS / Etc. Try it, and do your best work all while being away from that keyboard. Time to have fun!"

Sources1

OpenAI 把语音功能带进了 ChatGPT 桌面应用,Codex 负责人 Thibault Sottiaux 把它形容为科幻助手终于成真:"Jarvis / Samantha / TARS 等等。试试看,离开键盘也能完成你最好的工作。到了享受乐趣的时候了!"

AI educator Peter Yang

Peter Yang spent the day putting ChatGPT Voice through its paces, posting a before-and-after demo of his workflow.

Sources1

Peter Yang 花了一天时间深度体验 ChatGPT Voice,并发布了他工作流的前后对比演示。

His read on where this is heading: "The next evolution is being able to spin up multiple ChatGPT Voice threads so I can have a full team talking to me and to each other."

Sources1

他对这个方向的判断是:"下一步进化,是能同时开多个 ChatGPT Voice 线程,让一整个团队既跟我对话,也互相对话。"

His early feedback: it should notify you when other threads finish working, and the Chinese pronunciation "sounds bad."

Sources1

他的初步反馈:多线程并行时应该在其他线程完成后主动提醒,另外中文发音 "听起来很糟"。

Meta Sr Director of AI Madhu Guru

Following the GPT Sol incident, Madhu Guru shared takeaways from a chat with a security lead at a public company: identity and access management was designed for a finite number of employees, but one employee can now spin up hundreds of agents, and those agents can spawn more agents. The open questions: do agents inherit the spawning employee's permissions, what is an agent's lifecycle (a task, a ticket, a week?), do child agents inherit the same permissions, and how do you audit any of it?

Sources1

在 GPT Sol 事件之后,Madhu Guru 分享了他与一家上市公司安全负责人交流的心得:身份与访问管理(IAM)是为数量有限的员工设计的,但现在一个员工可以启动数百个 agent,这些 agent 还能再派生更多 agent。悬而未决的问题包括:agent 是否继承创建者的权限?agent 的生命周期是什么,一个任务、一张工单还是一周?子 agent 是否继承同样的权限?这一切又该如何审计?

Replit CEO Amjad Masad

Amjad Masad announced that the cost of Autoscale deployments, typically the most expensive part of running a scaled app, is down 80%.

Sources1

Amjad Masad 宣布 Autoscale 部署的费用下降了 80%,而这通常是规模化应用中最贵的一项开支。

He also showed off his chess autoresearch agent, which he says "got a PhD in modern LLM finetuning."

Sources1

他还展示了自己的国际象棋自动研究 agent,用他的话说,这个 agent "读完了一个现代 LLM 微调的博士学位"。

And he told the story of Viktor, a user who made money disrupting the agency model by building with Replit, then decided to automate the entire agency rather than just the coding part. Viktor asked Replit's team for an MCP, they built one, and he now runs an autonomous agency. As Masad puts it: "The Agency: It's merely an agent loop..."

Sources1

他还讲了用户 Viktor 的故事:Viktor 先是用 Replit 颠覆传统 agency 模式赚到了钱,随后决定干脆把整个 agency 自动化,而不只是写代码这一环。他向 Replit 团队要一个 MCP,团队做了一个,现在他已经跑起了自己的自主 agency。用 Masad 的话说:"所谓 Agency,不过是一个 agent 循环……"

Vercel CEO Guillermo Rauch

Guillermo Rauch announced that Python code now starts 2x faster on Vercel, automatically.

Sources1

Guillermo Rauch 宣布 Python 代码在 Vercel 上的启动速度快了一倍,而且是自动生效。

He also gave a shout-out to the AI Gateway team's shipping pace: "AI Gateway keeps getting better. Unreal product velocity from the team."

Sources1

他还称赞了 AI Gateway 团队的发布节奏:"AI Gateway 一直在变好。团队的产品迭代速度不可思议。"

Box CEO Aaron Levie

Aaron Levie argues the best way to think about AI is as a force multiplier for fields you already know, or for the rate at which you want to learn a new one. The third category, no existing judgment and no interest in developing it, "will basically produce slop." Experts, meanwhile, get better and better at their craft: they can steer agents properly, catch them when they veer off, and turn the output into something genuinely useful. His conclusion: specialization becomes more important as the tools get more powerful, because market expectations rise with them. "Getting good at any craft will continue to be necessary in the future."

Sources1

Aaron Levie 认为,理解 AI 的最佳方式是把它看作你已有专业领域的放大器,或者你学习新领域的加速器。至于第三类人,既没有已有的判断力、也无意培养判断力,"基本上只会产出 slop"。而专家会在自己的手艺上越来越强:他们知道如何正确引导 agent、在 agent 跑偏时把它拉回来,并把产出真正变成有用的东西。他的结论是:工具越强大,专业化反而越重要,因为市场的期望值也随之水涨船高。"把一门手艺练好,在未来依然是必需的。"

Y Combinator CEO Garry Tan

Garry Tan kept beating the open-source drum: "Open weight models are very very important."

Sources1

Garry Tan 继续为开源阵营发声:"开放权重模型非常非常重要。"

FPV Ventures partner Nikunj Kothari

Nikunj Kothari listed the tech titles that have "lost all signal" from overuse: "neo"-anything, full stack, fellows, labs, partner, forward deployed, and RL ("getting there slowly"), while owning the irony that he runs a fellowship and his own title is partner.

Sources1

Nikunj Kothari 列出了那些因为被滥用而 "彻底失去信号" 的科技圈用语:"neo" 开头的一切、full stack、fellows、labs、partner、forward deployed,还有正在沦陷的 RL。他也自嘲了其中的讽刺:自己既运营着一个 fellowship,头衔还是 partner。

Claude (Anthropic)

Anthropic upgraded Claude's voice mode: it now runs on the more capable models you have in chat, including Claude Opus and Sonnet, and can reach your connected tools mid-conversation, like email and calendar. Voice mode also supports more languages on every plan, including Spanish, French, Hindi, and Japanese. The update is rolling out in public beta on mobile, desktop, and web.

Sources1

Anthropic 升级了 Claude 的语音模式:语音对话现在使用你在聊天中可用的更强模型,包括 Claude Opus 和 Sonnet,还能在对话中途调用你已连接的工具,比如邮箱和日历。语音模式还在所有套餐中支持了更多语言,包括西班牙语、法语、印地语和日语。此次更新正以公开 beta 形式在移动端、桌面端和网页端推送。

OFFICIAL BLOGS

Claude Blog — Claude Code now supports artifacts

Claude Code can now capture work as artifacts: live, shareable visual pages built from the full context of a session, including your codebase, your connectors, and the conversation itself. Think PR walkthroughs, system explainers, filterable dashboards, incident timelines, and release checklists that fill themselves out as work gets done. "With artifacts, you don't need to wire up data sources or stand up infrastructure. You ask for a page, and Claude Code builds it from what already exists." Every publish is a new version at the same link, the open page refreshes in place for everyone watching, and version history lets you restore any time. Anthropic says its most common internal use case has been debugging: an engineer kicks off an incident investigation, Claude Code publishes a timeline with suspect commits and an error-rate chart, then republishes as the investigation progresses, so nobody has to "walk us through what the agent found." Artifacts are private to their author by default and shareable only within an authenticated org (they cannot be made public), with admin toggles, role-based scoping, retention policies, and a compliance API. Available in beta for Claude Team and Enterprise orgs, from the Claude Code CLI and desktop app.

Sources1

Claude Code 现在可以把工作进展沉淀为 artifact:基于会话完整上下文(代码库、connector 和对话本身)生成的实时、可分享的可视化页面。典型形态包括 PR 讲解、系统架构说明、可筛选的 dashboard、事故时间线,以及随工作推进自动填写的发布清单。"有了 artifacts,你不需要接数据源,也不需要搭基础设施。你只要开口要一个页面,Claude Code 就会用已有的一切把它做出来。" 每次发布都是同一链接下的新版本,打开的页面会原地刷新,所有人看到的都是最新内容,版本历史随时可以回滚。Anthropic 表示内部最常见的用法是 debug:工程师发起一次事故调查,Claude Code 发布一个包含时间线、可疑 commit 和错误率图表的 artifact,并随着调查推进不断重新发布,团队不再需要 "让人复述一遍 agent 发现了什么"。Artifact 默认仅作者可见,只能在完成认证的组织内部分享(无法公开),并配有管理员开关、基于角色的权限、保留策略和合规 API。目前以 beta 形式向 Claude Team 和 Enterprise 组织开放,可从 Claude Code CLI 和桌面应用使用。

PODCASTS

The MAD Podcast with Matt Turck — The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The Takeaway: Inference speed has become the AI industry's true bottleneck, and the supply shortages everyone fears most (HBM memory, advanced packaging, 3nm fabs) are exactly the things Cerebras' decade-old wafer-scale bet was built to sidestep.

核心要点: 推理速度已经成为 AI 行业真正的瓶颈,而所有人最担心的供应短缺(HBM 内存、先进封装、3 纳米产能)恰恰是 Cerebras 十年前那场晶圆级豪赌天然绕开的东西。

Andrew Feldman is cofounder and CEO of Cerebras, the company behind the largest chip in computing history: a wafer-scale processor 58 times the size of a GPU, stuffed with ultra-fast SRAM. Cerebras just pulled off the biggest semiconductor IPO of all time and signed a deal north of $20 billion to supply OpenAI with 750 megawatts of inference capacity through 2028.

Andrew Feldman 是 Cerebras 的联合创始人兼 CEO,这家公司造出了计算史上最大的芯片:一块比 GPU 大 58 倍、塞满超高速 SRAM 的晶圆级处理器。Cerebras 刚完成了半导体史上最大的 IPO,并与 OpenAI 签下了超过 200 亿美元的合同,将在 2028 年前为其提供 750 兆瓦的推理算力。

For years, AI was "a parlor trick." Around mid-2025, models got smart enough that people actually started using them, and the moment AI became productive, the metric that matters became tokens per second per user. "How big is the market for slow search? How big is the market for dial-up? It's zero." His Netflix analogy: when the internet got fast, Netflix didn't get more efficient at mailing DVDs, it became a movie studio. Speed doesn't just make AI faster, it changes what AI can be.

他说 AI 曾经长期只是个 "杂耍把戏",直到 2025 年年中,模型聪明到人们真正开始用它。而一旦 AI 变得有生产力,最重要的指标就成了每用户每秒 token 数。"慢速搜索的市场有多大?拨号上网的市场有多大?是零。" 他有个 Netflix 类比:互联网变快之后,Netflix 没有去优化邮寄 DVD 的效率,而是变成了一家电影公司。速度不只是让 AI 更快,它会改变 AI 本身能成为什么。

Why GPUs struggle with inference comes down to physics: generating each word requires moving all of a model's weights from memory to compute. For a 70-billion-parameter model, that's the equivalent of moving 100 HD movies of data per word, repeated a thousand times for a thousand-word answer. HBM memory is too slow for that. SRAM is blisteringly fast but can't store much, so Cerebras built a dinner-plate-sized chip and packed it with SRAM, moving weights roughly 2,500 times faster than a GPU.

GPU 做不好推理,根源在物理层面:每生成一个词,都要把模型的全部权重从内存搬到计算单元。对一个 700 亿参数的模型来说,这相当于每个词都要搬运 100 部高清电影的数据量,一条千词回答要重复搬一千次。HBM 内存对此太慢了,SRAM 快得惊人但存不了多少,于是 Cerebras 造了一块餐盘大小的芯片,把它塞满 SRAM,权重搬运速度大约是 GPU 的 2500 倍。

The industry's three bottlenecks right now: HBM, made by only three companies and sold out; TSMC's CoWoS advanced packaging, sold out; and 3nm fab capacity, heavily congested. Cerebras touches none of them, running on SRAM, its own packaging, and 5nm. He also flags a bottleneck nobody talks about: agentic AI is driving a CPU shortage, because "the AI processor is like the brains, and the CPUs are like the body," executing every website visit and data fetch an agent orders up.

行业当下的三大瓶颈:HBM 内存全球只有三家厂商生产且全部售罄,台积电的 CoWoS 先进封装产能售罄,3 纳米产能严重拥挤。Cerebras 一个都不碰:它用 SRAM、自研封装和 5 纳米工艺。他还点出一个没人讨论的瓶颈:agentic AI 正在推高 CPU 需求,因为 "AI 处理器是大脑,CPU 是身体",agent 每一次访问网页、抓取数据的动作都由 CPU 执行。

On moats, he's blunt: two years ago every state-of-the-art model was trained in a CUDA flow; today Gemini trains on TPUs and Claude on Trainium, and he says CUDA has lost roughly 70% share of frontier training in a year or two. "There's no moat in inference. It takes you eight keystrokes to move from a GPU to us in the cloud." And Cerebras' own story is a decade in the desert: the team spent eighteen months burning $8 million a month to solve wafer-scale packaging, a problem nobody in computing history had solved, delivered it in 2020, and "nobody cared. Nobody bought it." AI was still a hobby, and nobody cares how fast their hobby is. The first generation barely sold, the second sold a few hundred, the third sold tens of thousands. "The way you have perfect timing is to have horrible timing for ten years."

关于护城河,他的判断很直接:两年前所有前沿模型都在 CUDA 体系里训练,如今 Gemini 用 TPU 训练、Claude 用 Trainium 训练,他说一两年间 CUDA 在前沿训练中丢掉了大约七成份额。"推理没有护城河。从 GPU 切到我们的云,只需要敲八个键。" 而 Cerebras 自己的故事是十年冷板凳:团队花了十八个月、每月烧 800 万美元,解决了计算史上没人解决过的晶圆级封装问题,2020 年交付时却 "没人在乎,没人买",因为当时 AI 只是个爱好,没人在乎自己的爱好快不快。第一代芯片几乎没卖出去,第二代卖了三五百台,第三代卖出数万台。"获得完美时机的方式,就是先经历十年糟糕的时机。"

Sources1
Generated through the Follow Builders skill — bilingual daily signal and weekly perspective from the people building AI.