X / Twitter
Swyx (swyx on X) — Latent Space co-host, Cognition Swyx shared his current multi-model stack for serious ("Big Boy") projects: Sol Ultra to plan, Fable 5 to critique, Sonnet 5 / Terra Ultra / SWE 1.7 to write code, and Devin review to review. His meta-tip: always run a "grill-me" or "interview-me" style prompt first to force the key decisions upfront before any code gets written.
Swyx 分享了他做重要("Big Boy")项目时的多模型组合:用 Sol Ultra 做规划,用 Fable 5 做批判审视,用 Sonnet 5 / Terra Ultra / SWE 1.7 写代码,再用 Devin review 做代码审查。他的方法论心得是:动手写代码前,先跑一个 "grill-me" 或 "interview-me" 风格的 prompt,把关键决策逼出来定下来。
Thibault Sottiaux (thsottiaux on X) — Codex & ChatGPT, OpenAI OpenAI's Codex and ChatGPT lead Thibault Sottiaux announced "ChatGPT Work" and teased that the team may be about to celebrate 8M active users.
OpenAI Codex 与 ChatGPT 负责人 Thibault Sottiaux 官宣了 "ChatGPT Work",并放话说团队可能马上要庆祝 8M 活跃用户了。
Thariq (trq212 on X) — Claude Code, Anthropic Anthropic's Thariq detailed the new Claude Artifacts upgrade, which makes artifacts much more expressive and composable. His favorite use: build a project dashboard in Claude Tag that others — or his own local Claude Code sessions — can edit.
Anthropic 的 Thariq 详细介绍了 Claude Artifacts 的最新升级,让 artifacts 的表达力和可组合性大大增强。他最喜欢的玩法:在 Claude Tag 里搭一个项目 dashboard,让别人、或者他自己本地的 Claude Code 会话都能编辑它。
Amjad Masad (amasad on X) — CEO, Replit Replit CEO Amjad Masad is now getting realtime progress updates on his model training runs, which he says "feels like early vibe coding except it's making personal models."
Replit CEO Amjad Masad 现在可以拿到自己模型训练过程的实时进度更新。他说这"感觉就像早期的 vibe coding,只不过这次做的是个人专属模型"。
Guillermo Rauch (rauchg on X) — CEO, Vercel Vercel CEO Guillermo Rauch noted that open-weight models now run 29% of AI Gateway tokens, up from 11% in April. He also highlighted giving agents the ability to set up and tune feature-flag experiments — a building block for autonomous, self-optimizing websites and applications.
Vercel CEO Guillermo Rauch 指出,open-weight 模型现在已占 AI Gateway token 用量的 29%,而 4 月时还只有 11%。他还着重提到,让 agent 具备自行搭建和调优 feature-flag 实验的能力,是构建"自主、自我优化的网站与应用"的关键基石。
Aaron Levie (levie on X) — CEO, Box Box CEO Aaron Levie laid out a structural map of AI's future: frontier intelligence keeps advancing while per-task pricing falls; open weights rapidly absorb frontier breakthroughs and can be run "at cost" and post-trained for specific domains; and the applied-AI layer wins by orchestrating frontier plus cheaper models with deep domain context. He argues model routing — frontier intelligence as the "manager," cheaper models as the workhorse — is the future template, and warns that training a model per enterprise "is going to be a lot harder than it looks" because your most valuable data is also your most sensitive and is constantly changing.
Box CEO Aaron Levie 勾勒了一张 AI 未来的结构图:frontier 智能持续进步,同时单任务价格不断下降;open weights 会迅速吸收前沿突破,既能在云上"按成本价"运行,又能针对特定领域做 post-training;而 applied-AI 这一层的赢法,是用深厚的领域上下文把 frontier 模型和更便宜的模型编排到一起。他认为 model routing——让 frontier 智能当"经理"、便宜模型当"主力干活的"——就是未来的范式模板;并提醒,"给每家企业单独训一个模型"会"比看上去难得多",因为你最有价值的数据往往也是最敏感、且一直在变的。
Ryo Lu (ryolu_ on X) — Design, Cursor Cursor designer Ryo Lu built custom e-reader firmware with Cursor — beautiful Latin + CJK typography, vertical layout with proper line-breaking (縱書.禁則), large character sets, book and progress sync with ryOS, and speedy rendering with caching. A fun proof that coding agents can now reach all the way down into firmware and hardware hacking.
Cursor 设计师 Ryo Lu 用 Cursor 给电子阅读器写了一套自定义固件——漂亮的拉丁 + 中日韩字体排版、纵向排版与正确的断行处理(縱書.禁則)、大字符集支持、书籍与阅读进度同步到 ryOS,还有带缓存的高速渲染。这有趣地证明了:编程 agent 现在已经能一路下探到固件层、玩起硬件折腾了。
Zara Zhang (zarazhangrui on X) — Builder Zara Zhang shared a framework for the 3 levels of AI adoption inside organizations, noting that most companies are still stuck at level 2.
Zara Zhang 分享了一个"组织内 AI 采用的三个层级"框架,并指出大多数公司都还卡在第 2 级。
Nikunj Kothari (nikunj on X) — Partner, FPV Ventures FPV Ventures partner Nikunj Kothari built and open-sourced a "Ramp-Autofill" skill using the Ramp CLI and Claude Fable: it finds receipts automatically from your iMessage and Gmail (using Playwright to convert linked web pages into PDFs), fills in memos from your Google Calendar events, learns your categorization style from past transactions, verifies its own work, and can run as a scheduled job. He used it to clear 60 days of expenses over a weekend, and noted one version was one-shot by Fable via a voice prompt while driving to SF.
FPV Ventures 合伙人 Nikunj Kothari 用 Ramp CLI 加 Claude Fable 做了一个开源的 "Ramp-Autofill" skill:它能自动从你的 iMessage 和 Gmail 里找收据(用 Playwright 把链接里的网页转成 PDF),根据你的 Google Calendar 日程自动填写备注,从历史交易里学会你的分类习惯,还会自我校验、并可作为定时任务运行。他用它一个周末就清掉了过去 60 天的报销,还提到其中一个版本是 Fable 在他开车去 SF 的路上、用语音 prompt 一次成型做出来的。
Peter Steinberger (steipete on X) — OpenClaw, OpenAI Peter Steinberger shipped updates to OpenClaw's iOS and Android apps (with a Node version bump for stability) and shared a handy prompt tip: "stress test" is a good prompt.
Peter Steinberger 发布了 OpenClaw 的 iOS 和 Android app 更新(顺带升级了 Node 版本以提升稳定性),并分享了一个实用的 prompt 小技巧:"stress test"(压力测试)本身就是个好 prompt。
Sam Altman (sama on X) — CEO, OpenAI Reflecting on model progress, OpenAI CEO Sam Altman said it "still sorta breaks my brain to see our models be good at design finally."
谈到模型的进步时,OpenAI CEO Sam Altman 说,"看到我们的模型终于擅长做设计了,还是有点让我脑子转不过来。"
Official Blogs
Claude Blog
Building intelligent apps for Apple platforms with Claude in the Foundation Models framework Anthropic released Foundation Models framework support for Claude via a new Swift package, letting Apple developers call Claude from Apple's own Foundation Models framework for more complex workflows. Apple's framework returns typed Swift values through guided generation (via @Generable annotations) in as few as three lines of code, so developers can lean on fast on-device models for local tasks like summarization or extraction, then hand off to Claude when a request needs multi-step reasoning, code generation, web search, or code execution. The examples make the pattern concrete: a journaling app generates daily prompts on-device, then asks Claude to find threads across months of entries; a study app defines a term locally, then hands off to Claude when the student asks "why does this matter for everything else we've covered?" As Anthropic puts it, "It's one experience for the user, backed by the right model for each step." Support arrives on iOS 27, iPadOS 27, macOS 27, visionOS 27, and watchOS 27 — add the package, sign in with an Anthropic API key, and pass typed outputs from Apple's on-device pass into a Claude request; the package handles streaming, tool calls, and structured responses back into your SwiftUI view.
Anthropic 通过一个新的 Swift package 发布了对 Claude 的 Foundation Models framework 支持,让 Apple 开发者可以直接从 Apple 自家的 Foundation Models framework 里调用 Claude 来处理更复杂的工作流。Apple 的这套 framework 通过 guided generation(借助 @Generable 注解)能用短短三行代码返回带类型的 Swift 值,因此开发者可以用快速的端侧模型处理摘要、信息抽取这类本地任务,再在需要多步推理、代码生成、联网搜索或执行代码时把请求交接给 Claude。文中的例子把这套模式讲得很具体:一个日记 app 在端侧生成每日 prompt,再让 Claude 在数月的记录里找出贯穿的线索;一个学习 app 在本地给出术语定义,等学生追问"这跟我们学过的其他东西有什么关系?"时再交接给 Claude。用 Anthropic 的话说,"对用户来说是一段连贯的体验,但每一步背后都由最合适的模型来支撑。" 该支持将登陆 iOS 27、iPadOS 27、macOS 27、visionOS 27 和 watchOS 27——只需加入 package、用 Anthropic API key 登录,把 Apple 端侧那一步产出的带类型输出传进一个 Claude 请求即可;package 会处理流式返回、工具调用,以及把结构化结果送回你的 SwiftUI 视图。
Podcasts
The MAD Podcast with Matt Turck — Inside Nemotron & NVIDIA's AI Lab | Bryan Catanzaro
The Takeaway: When compute is capped — by power, dollars, or silicon — the only way to get more intelligence is to be more efficient, not to apply more force.
Bryan Catanzaro leads Nemotron, NVIDIA's family of open foundation models, and has been building AI at NVIDIA since 2008 — he created cuDNN and the Megatron project — with a detour to Baidu's Silicon Valley AI Lab alongside Andrew Ng and a young Dario Amodei. His central claim reframes the entire scaling debate: "if you accept as the truth that we're gonna be running at the limit, then... the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit." Moore's Law, he insists, has been dead for years, so progress now comes from specialization and co-design — understanding AI deeply enough to shape the chips themselves.
That philosophy explains Nemotron's engineering choices: pretraining in 4-bit (NVFP4) arithmetic, a hybrid architecture that is mostly state-space model with a little attention (which he argues is not just faster but genuinely smarter), mixture-of-experts with a "latent MoE" trick that quadruples the number of experts at the same inference cost, multi-token prediction that makes a more accurate model also a faster one, and a 1M-token context window.
His most refreshing takes are human, not technical. He flatly rejects the "China just copies via distillation" narrative — having worked at Baidu, he calls it "absolutely false" — and frames multi-teacher distillation as an org-design tool: "one of the challenges of building AI in 2026 is that you have to figure out how to get the people to work together even though you're only building one thing at the end of the day." Nemotron, he says, exists both to keep NVIDIA sharp enough to design the next system and to keep open AI thriving — because whenever AI gets deployed, it's good for NVIDIA's business.
一句话要点:当算力被卡死时——不管是被电力、被预算还是被硅片本身卡住——想要更多智能的唯一办法就是变得更高效,而不是砸更多蛮力。
Bryan Catanzaro 领导着 Nemotron——NVIDIA 的开源基础模型家族。他从 2008 年就在 NVIDIA 做 AI,创造了 cuDNN 和 Megatron 项目,中间还去百度硅谷 AI 实验室待过一段,同事里有 Andrew Ng,还有年轻时的 Dario Amodei。他的核心论断重新定义了整场 scaling 之争:"如果你承认一个事实:我们注定会一直贴着极限在跑,那么……获得更多智能的办法就是变得更高效。既然已经到极限了,就不可能靠砸更多蛮力换来更多智能。" 他坚称摩尔定律已经死了好些年,所以如今的进步只能来自专用化与协同设计——把 AI 理解得足够深,深到能反过来塑造芯片本身。
这套理念解释了 Nemotron 的一系列工程选择:用 4-bit(NVFP4)精度做预训练;采用一个以 state-space model 为主、只掺一点点 attention 的混合架构(他认为这不只是更快,而是真的更聪明);用带 "latent MoE" 技巧的 mixture-of-experts,在相同推理成本下把专家数量翻两番;用 multi-token prediction 让"更准的模型同时也更快";以及一个 1M token 的上下文窗口。
而他最令人耳目一新的观点其实关乎人、而非技术。他直接否定了"中国只是靠蒸馏抄袭"这套说法——因为他在百度工作过,直言这"绝对是错的"——并把 multi-teacher distillation 看作一种组织设计工具:"2026 年做 AI 的一个挑战是,你得想办法让一群人协作,可到头来大家其实只在造同一个东西。" 他说,Nemotron 存在的意义有两层:既让 NVIDIA 保持足够敏锐、去设计下一代系统,又让开源 AI 持续繁荣——因为只要 AI 被部署出去,就对 NVIDIA 的生意有好处。