回到卷首
每日集录ai builders

六月二十七日

二〇二六年 15 builders 34 posts 1 podcast 约三十二分钟

OpenAI shipped GPT-5.6 'Sol' but a U.S. government directive caps initial access to ~20 pre-approved companies, drawing fire from Y Combinator's Garry Tan and Every's Dan Shipper while Box's Aaron Levie calls the model very strong with 'no walls' in sight; Sam Altman teased an updated GPT-5.5 instant model and OpenAI reset all Codex users' usage; Peter Yang warned gating frontier models just makes open source more attractive; Vercel's Guillermo Rauch dissected why agents are uniquely hard to debug and crowned shadcn 'the UI for AI'; FPV's Nikunj Kothari argued taste only comes from being in the arena; and on No Priors, OpenAI's Noam Brown explained why benchmark grids mislead once you fail to control for test-time compute.

The big story today: OpenAI shipped GPT-5.6 "Sol," but a U.S. government directive initially limits access to roughly 20 pre-approved companies — igniting a debate among builders about gated frontier models, while No Priors digs into why benchmark grids are quietly broken.

今日头条:OpenAI 发布了 GPT-5.6 “Sol”,但一纸美国政府指令把初期访问权限制在约 20 家预先批准的公司手中,由此引爆了一场关于“封闭式前沿模型”的争论;与此同时,No Priors 深入剖析了为什么各家发布的 benchmark 网格其实早已失真。

X / TWITTER

Every CEO Dan Shipper broke the news that OpenAI's new GPT-5.6 "Sol" is, by U.S. government directive, initially limited to roughly 20 pre-approved companies — and Every is not on the list. He framed it as likely temporary while Washington works out a policy for releasing frontier models, saying he "applauds the need for some government oversight" of cyber-resilient infrastructure but warns that locking advanced models to a handful of AI giants would deny "ambitious students, independent builders, and working professionals the tools they need to learn, create, and compete." His core worry: early testers like Every exist to prepare Americans to use these tools well, and losing access guts that mission.

Sources1

Every CEO Dan Shipper 率先爆料:根据一项美国政府指令,OpenAI 新发布的 GPT-5.6 “Sol” 初期只对约 20 家预先批准的公司开放,而 Every 不在名单上。他认为这很可能是华盛顿在敲定前沿模型发布政策期间的临时措施,并表示自己“理解并支持政府对网络韧性基础设施进行一定程度的监管”,但警告说,把先进模型锁给少数几家 AI 巨头,会让“有抱负的学生、独立开发者和职场专业人士”失去“学习、创造和竞争所需要的工具”。他最担心的是:像 Every 这样的早期测试者本就是为了帮美国人学会用好这些工具而存在,一旦失去访问权,这个使命也就无从谈起。

Y Combinator President and CEO Garry Tan sharply criticized the gated rollout, warning that "this is honestly no way to release a model" and that continuing to ship frontier models this way is "a solid way to salt the ground and kill all innovation by small startups." The take crystallizes founder anxiety that restricting the best models structurally disadvantages the startups YC backs.

Sources1

Y Combinator 总裁兼 CEO Garry Tan 对这种封闭式发布提出尖锐批评,直言“说实话,没有谁该这样发布一个模型”,并称继续以这种方式推出前沿模型“是一种把土地撒盐、扼杀所有小型创业公司创新的稳妥做法”。这番话道出了创业者们的普遍焦虑:把最强模型设限,会从结构上让 YC 所投的初创公司处于劣势。

Box CEO Aaron Levie offered the bullish counterpoint on the model itself: "GPT-5.6 is real and looks very strong," especially for knowledge-worker tasks that lean on heavy tool use and long-running agents. His headline claim: "We're not hitting any walls in AI progress right now."

Sources1

Box CEO Aaron Levie 则给出了对模型本身的乐观判断:“GPT-5.6 是真的,而且看起来非常强”,尤其擅长那些重度依赖工具调用和长时间运行 agent 的知识工作类任务。他的核心论断是:“目前 AI 的进展并没有撞上任何天花板。”

AI educator Peter Yang zeroed in on the strategic absurdity of gating frontier models: the U.S. publishes frontier models, they get distilled into cheap open-source models, U.S. companies adopt those open-source models because they're "good enough and much cheaper," and then access to frontier models gets gated — asking whether the net effect is simply that U.S. companies innovate less while open source becomes more attractive. Separately, he argued the money has moved to services (with some software bundled) rather than pure software, because "people want outcomes, not tools," making it hard to build a pure-play software company more valuable than Codex or Claude Code wired up with personal skills and agents. He also posted a Claude Code wishlist: steer conversations mid-task, mobile remote control on by default, better keyboard shortcuts, and drag-to-reorder projects.

Sources123

AI 科普作者 Peter Yang 抓住了封锁前沿模型在战略上的荒诞之处:美国发布前沿模型 → 它们被蒸馏成廉价的开源模型 → 美国公司因为这些开源模型“够用且便宜得多”而纷纷采用 → 然后前沿模型的访问权又被设限。他追问:这么折腾下来,结果是不是只是让美国公司创新更少、而开源反倒更有吸引力?另外,他指出钱已经从纯软件流向了服务(外加打包一点软件),因为“人们要的是结果,而不是工具”,这让人很难再做出一家比“Codex 或 Claude Code + 一堆个人 skills 和 agent”更有价值的纯软件公司。他还列了一份 Claude Code 愿望清单:希望能在任务执行过程中插话引导、默认开启移动端远程控制、改进快捷键,以及支持拖拽重排项目。

OpenAI CEO Sam Altman teased that the team updated the GPT-5.5 instant model used in ChatGPT this week ("i like its vibes") and hinted at pricing direction, saying it's "not quite all-you-can-eat tokens, but we are working on it."

Sources12

OpenAI CEO Sam Altman 透露团队本周更新了 ChatGPT 里使用的 GPT-5.5 instant 模型(“我喜欢它的感觉”),还暗示了定价方向,说目前“还做不到 token 管够随便用,但我们正在努力。”

OpenAI's Thibault Sottiaux, who works on Codex and ChatGPT, announced OpenAI is giving all Codex users a usage reset "on the house," appearing in accounts within hours. He added that OpenAI applied mitigations and its investigation "hasn't shown users being impacted at large," but is continuing to monitor.

Sources1

负责 Codex 与 ChatGPT 的 OpenAI 工程师 Thibault Sottiaux 宣布,OpenAI 将免费为所有 Codex 用户重置用量额度,几小时内就会在账户里生效。他补充说,OpenAI 已经采取了一些缓解措施,调查“并未显示用户受到大范围影响”,但会继续密切监控。

Vercel CEO Guillermo Rauch dug into why agents are uniquely hard to debug: by design AI models are non-deterministic (even identical prompts can diverge), and agents are also complex distributed systems spanning many functions, sandboxes, and dozens of API services that can fail or rate-limit you. He said baking in out-of-the-box observability was a key priority for Vercel's agent product. Separately, he declared "The UI for AI is here. It's shadcn," endorsing the component library as the default interface layer for AI apps.

Sources12

Vercel CEO Guillermo Rauch 深入剖析了 agent 为何格外难调试:AI 模型在设计上就是非确定性的(哪怕完全相同的 prompt,输出也可能不一样),而 agent 同时又是复杂的分布式系统,横跨众多函数、sandbox,以及几十个随时可能宕机或限流的 API 服务。他说,把开箱即用的可观测性内建进来,是 Vercel agent 产品的重点之一。另外,他还断言“给 AI 用的 UI 已经到来了,那就是 shadcn”,力挺这个组件库作为 AI 应用默认的界面层。

FPV Ventures partner Nikunj Kothari pushed back on the wave of "taste" commentary, arguing you can't acquire or refine taste without being in the arena: like a chef whose 101st dish carries the lessons of 10,000 iterations, taste comes from doing, not spectating. His twist on the AI angle: over-fitting without variation makes work stale, the best people innovate by consistently breaking patterns, and — unlike most skeptics — he thinks AI "has a fair shot at actually building taste."

Sources1

FPV Ventures 合伙人 Nikunj Kothari 对当下铺天盖地的“品味(taste)”论调提出反驳:他认为不亲自下场,就无法获得或打磨品味——就像一位厨师,第 101 道菜里凝结的是前 10,000 次迭代的经验,品味来自动手做,而非旁观。落到 AI 上他还补了一刀:缺乏变化的过度拟合会让作品变陈旧,真正顶尖的人靠持续打破套路来创新;而与大多数怀疑者不同,他认为 AI“确实有机会真正建立起品味”。

Linear head of product Nan Yu shared a sharp product-judgment heuristic he calls "Secret level 6: there's a problem, but it's not worth solving, so leave it alone." His point: organizations full of "level 1s" can win precisely because they don't get distracted chasing every side quest.

Sources1

Linear 产品负责人 Nan Yu 分享了一条犀利的产品判断法则,他称之为“隐藏第 6 级:问题确实存在,但不值得解决,那就别去碰它。”他的意思是:满是“1 级选手”的组织之所以能赢,恰恰是因为他们不会被各种支线任务分散注意力。

Builder Zara Zhang recommended Borumi as the most underrated screen-recording and video-editing tool she's found — "like Screen Studio + Descript + CapCut all in one" — noting she used it to make her latest video.

Sources1

独立开发者 Zara Zhang 推荐了 Borumi,称它是她用过的最被低估的录屏与视频剪辑工具——“相当于把 Screen Studio + Descript + CapCut 合三为一”——并提到自己最新的视频就是用它做的。

South Park Commons general partner Aditya Agarwal shared a reflection on AI's social side effect: it's left him with "zero tolerance for shallow interactions with humans," pushing him to value depth in his relationships while delegating everything transactional to agents. His prediction: "our world will become both smaller and more rich in its relationship depth."

Sources1

South Park Commons 普通合伙人 Aditya Agarwal 分享了他对 AI 一个社会副作用的思考:它让他对“与人之间浅层的互动几乎零容忍”,反倒促使他更看重关系中的深度,把一切事务性的事都交给 agent。他的预测是:“我们的世界会变得既更小、又在关系深度上更丰富。”

Latent Space cohost and AI Engineer organizer Swyx flagged a hiring-market signal: with both OpenAI and Anthropic standing up multi-billion-dollar services arms, the Forward Deployed Engineer (FDE) is now "one of the most in-demand disciplines on earth." He's running a first-of-its-kind AI FDE miniconference to build out coverage of the role.

Sources1

Latent Space 联合主持人、AI Engineer 大会组织者 Swyx 指出了一个招聘市场信号:随着 OpenAI 和 Anthropic 都在组建数十亿美元规模的服务(services)部门,Forward Deployed Engineer(FDE,前沿部署工程师)如今已是“地球上最抢手的职业之一”。他正在筹办一场首创的 AI FDE 小型大会,来补齐这一岗位的相关内容。

PODCASTS

No Priors — "Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown"

The Takeaway: A modern model's intelligence is now a function of how much money you spend at inference — and the benchmark grids everyone publishes hide exactly that, making real progress look smaller than it is.

Noam Brown, an OpenAI research scientist who pioneered AI reasoning and test-time compute (the work behind models that "think" before answering), argues the industry is measuring its own models wrong. When GPT-5.5 launched, the benchmark grid showed only a few points of improvement over 5.4, fueling skepticism — until people actually used it. The catch: those grids report a single number per model and never control for how much compute each model burned to get it. "Once you control for the amount of thinking time, actually, you can see that 5.5 is a substantial jump over 5.4." His fix is simple: put an x-axis on every benchmark — tokens, cost, or time — or fix a budget, instead of pretending one number tells the story.

Why it matters beyond marketing: modern models can now think productively for weeks, so the point where performance plateaus is "simply too far out to reasonably test." That breaks safety evaluation too. Preparedness frameworks were built in the ChatGPT era, when you couldn't usefully spend more compute at inference. Today, Brown warns, "the capability of the model is a function of how much money you put into it" — and dangerous capabilities scale with budget just like the useful ones.

The most striking example is latent capability already sitting inside shipped models. OpenAI disproved the Erdős unit-distance conjecture with an internal model, and people later coaxed the same result out of GPT-5.5 with the right scaffolding — for somewhere between $1,000 and $100,000 of compute. Nobody had simply asked: what happens if I pour $100k into 5.5? His most grounded take cuts against AI-doom narratives: there's no overnight intelligence explosion coming, because unlocking a model's best work still takes real wall-clock time. "The biggest bottleneck for all of us is time" — which is exactly why every frontier researcher is grinding right now.

Sources1

核心要点: 一个现代模型的智能水平,如今取决于你在推理(inference)阶段愿意花多少钱——而各家发布的 benchmark 网格恰恰把这一点藏了起来,让真实的进步看上去比实际小得多。

Noam Brown 是 OpenAI 的研究科学家,也是 AI 推理与 test-time compute(即让模型在回答前先“思考”的那套工作)的开创者之一。他认为整个行业都在用错误的方式衡量自己的模型。GPT-5.5 发布时,benchmark 网格显示它只比 5.4 高出区区几个百分点,引来一片质疑——直到人们真正上手用了它。问题在于:那些网格对每个模型只给一个数字,却从不控制各模型为拿到这个分数到底烧了多少算力。“一旦你把思考时长这个变量控制住,就能看出 5.5 相对 5.4 是一次实质性的跃升。”他给的解法很简单:给每个 benchmark 加上一条 x 轴——token、成本或时间——或者干脆设定一个预算,而不是假装单个数字就能说明全部。

为什么这件事远不止是营销话术:现代模型如今可以富有成效地“思考”数周之久,因此性能真正趋于平台的那个点“实在太远,根本无法合理地测出来”。这也让安全评估失灵了。各家的 preparedness framework 都诞生在 ChatGPT 时代,那会儿你在推理阶段多花算力也榨不出更多能力。而今天,Brown 警告说,“模型的能力,本质上是你往里投了多少钱的函数”——危险能力会随预算放大,和有用能力一样。

最惊人的例子,是那些早已藏在已发布模型里的潜在能力。OpenAI 用一个内部模型推翻了 Erdős 单位距离猜想;后来人们用合适的 scaffolding,从 GPT-5.5 里也“哄”出了同样的结果——所花算力大约在 1,000 到 100,000 美元之间。此前根本没人问过这样一个简单的问题:如果我往 5.5 里砸 10 万美元算力,会发生什么?他最务实的一个判断,正好戳破了 AI 末日论:不会有什么一夜之间的智能爆炸,因为要榨出一个模型的最佳表现,依然需要实打实的墙上时钟时间。“对我们所有人来说,最大的瓶颈就是时间”——这也正是为什么每一位前沿研究者现在都在拼命。

Sources1
Generated through the Follow Builders skill — bilingual daily signal and weekly perspective from the people building AI.