回到卷首
每日集录ai builders

八月七日

二〇二六年 15 builders 31 posts 1 podcast 1 blog 约三十三分钟

OpenAI makes GPT-5.6 Luna text chats unlimited for free users while improving Sol in chat; Anthropic cuts Fable 5 biology fallbacks by about 85% while retaining dual-use boundaries; Guillermo Rauch proposes a portable Plugin standard for agents; Aaron Levie frames enterprise adoption as governed workflow redesign; Amjad Masad argues coding models supersede no-code; Claude connects to Apple's Foundation Models framework; and Basis co-founder Mitch Troyanovsky explains why long-horizon agents need behavior specs, process supervision, and structured runtime context.

Top Signals

Free ChatGPT text moves to unlimited use as GPT-5.6 Sol improves in chat

OpenAI's Thibault Sottiaux says free ChatGPT users now have unlimited text chats powered by GPT-5.6 Luna. Sam Altman paired that distribution change with a model update, saying GPT-5.6 Sol is now much better in chat. Sottiaux also describes a sharp expansion in the practical scope of coding agents: after speaking to Codex with Sol for five minutes about work that appears to require weeks, he can step away and return to a completed result. Taken together, the posts point to a two-sided product strategy: make everyday text intelligence effectively unconstrained for free users while pushing the higher-end agent toward longer, more ambitious execution.

Sources123

OpenAI 的 Thibault Sottiaux 表示,ChatGPT 免费用户现在可以无限使用由 GPT-5.6 Luna 驱动的文本对话。Sam Altman 在宣布这一分发变化的同时也提到模型更新,称 GPT-5.6 Sol 的聊天表现已经明显提升。Sottiaux 还描述了 coding agent 实际工作范围的快速扩张:他用五分钟向搭载 Sol 的 Codex 讲述一项看起来需要数周的工作,短暂离开后回来,结果已经完成。把这些信息放在一起,可以看到一种双线产品策略:一边让免费用户几乎不受限制地使用日常文本智能,另一边则让高端 agent 承担更长、更有野心的执行任务。

Sources123

Claude Fable 5 cuts biology false positives while retaining a dual-use boundary

Anthropic says an update to Claude Fable 5's biology safeguards reduced biology-related fallbacks by about 85% across its product surfaces in testing. The change is intended to let Fable answer a wider range of ordinary health and educational questions, but it does not open unrestricted professional biology work: requests considered dual-use, including virology, toxicology, and molecular design, still fall back to Opus 5. Anthropic says Fable is therefore not yet usable for professional biology research and drug development, and frames trusted-access pathways as the route to closing that gap.

Sources1

Anthropic 表示,Claude Fable 5 的 biology safeguard 更新在测试中把各产品界面与生物学相关的 fallback 降低了约 85%。这项调整的目标是让 Fable 能回答更广泛的日常健康与教育问题,但并不意味着专业生物学能力已经完全开放:被认定具有 dual-use 风险的请求,包括病毒学、毒理学与分子设计,仍会 fallback 到 Opus 5。Anthropic 明确表示,Fable 目前仍不能用于专业生物学研究和药物开发,并把 trusted access pathway 视为未来缩小这一能力缺口的路径。

Sources1

The agent ecosystem is converging on open, portable extensions

Vercel CEO Guillermo Rauch argues that developer tools must be both open source and universally extensible, and calls AI coding agents the most important developer tools in the industry's history. His concrete proposal is a common Plugin standard: build an extension once and expose it across CLIs, IDEs, cloud agents, and personal assistants. The strategic consequence is larger than plug-in convenience. If software creation increasingly flows through many agent surfaces, a portable extension layer gives independent builders a way to reach that demand without rebuilding an integration for every host.

Sources1

Vercel CEO Guillermo Rauch 认为,developer tool 必须同时具备 open source 与普遍可扩展两项特征,并称 AI coding agent 是行业历史上最重要的开发者工具。他给出的具体方案是一套通用 Plugin 标准:扩展只需构建一次,就能覆盖 CLI、IDE、cloud agent 与 personal assistant。这一变化的战略意义不只是安装插件更方便。如果越来越多的软件生产经过不同的 agent 界面完成,那么可移植的扩展层就能让独立 builder 触达这股需求,而无需为每个宿主重复开发集成。

Sources1

Enterprise agents require workflow redesign, not chatbot-style prompting

Box CEO Aaron Levie says working with an agent resembles managing someone inside a process more than asking a chatbot a question. In his formulation, an agent prompt is closer to a specification: the task needs scope and an explicit definition of done. The larger payoff comes only when companies redesign the workflow itself, giving agents the right data, allowing work to cross organizational boundaries, and changing where humans review results. Levie also argues that systems of record can become more important as agents generate far more code and process more data, because enterprises still require governance, security, compliance, guardrails, and controlled access. His forecast is that most enterprise token usage will eventually come from deployed agents executing work inside these governed processes.

Sources12

Box CEO Aaron Levie 认为,与 agent 协作更像是在流程中管理一个人,而不是向 chatbot 提问。按他的说法,给 agent 的 prompt 更接近 specification:任务需要清晰限定范围,还要明确什么叫做完成。真正更大的收益只有在企业重新设计 workflow 后才会出现,包括向 agent 提供正确数据、允许工作跨越组织边界,以及改变人类审核结果的环节。Levie 同时指出,当 agent 生成更多代码并处理更多数据时,system of record 反而可能变得更重要,因为企业仍然需要 governance、安全、合规、guardrail 与受控的数据访问。他预测,企业中的大部分 token 用量最终会来自被部署到这些受治理流程中执行任务的 agent。

Sources12

Coding models are replacing the old boundary between code and no-code

Replit CEO Amjad Masad argues that Airtable marks both the rise and fall of the no-code era. His reasoning is that a graphical interface cannot express arbitrary software, so making software creation broadly accessible always required solving code itself. Masad adds historical context: in 2021 and 2022 he asked Google, Meta, and others to train coding-specific models, found little interest compared with NLP use cases, and ultimately trained Replit-code-3b. His retrospective claim is that the industry eventually became convinced of code as a primary model capability, validating an approach that previously sounded implausible.

Sources12

Replit CEO Amjad Masad 认为,Airtable 同时标志着 no-code 时代的兴起与结束。他的理由是,图形界面无法表达任意软件,因此要让软件开发真正普及,始终需要解决 code 本身。Masad 还补充了一段历史:2021 和 2022 年,他曾邀请 Google、Meta 等公司一起训练 coding-specific model,但当时与 NLP use case 相比几乎无人重视,Replit 最终自行训练了 Replit-code-3b。他的回顾性判断是,行业后来终于把 code 视为模型的核心能力,这也验证了一条当初听起来近乎不切实际的路线。

Sources12

Consumer AI advantage may come from onboarding and trust, not the best model

AI educator Peter Yang believes the consumer market is largely ChatGPT's and Google's to lose. ChatGPT's task is to turn its one billion users into people who connect apps and let agents perform work, while overcoming both distrust of broad data access and low awareness of current capabilities. Gemini, he argues, benefits from nearly one billion users and a familiar Google identity spanning email, calendar, Workspace, and Chrome, but trails in third-party plugins, browser use, and other features. His broader point is that ordinary users care less about which frontier model is best than whether onboarding is clear, pricing is fair, and the product completes work reliably. Marketing and product messaging therefore matter alongside model progress.

Sources1

AI 教育者 Peter Yang 认为,消费级 AI 市场很大程度上已经是 ChatGPT 和 Google 不容有失的阵地。ChatGPT 的任务,是让其十亿用户愿意连接常用 app,并授权 agent 执行实际工作;主要障碍既包括用户不信任 AI 获取广泛数据权限,也包括普通人不了解 ChatGPT 现在已经能做什么。他认为 Gemini 拥有接近十亿用户,以及覆盖 email、calendar、Workspace 和 Chrome 的熟悉 Google 身份体系,但在第三方 plugin、browser use 和其他关键功能上仍然落后。他更大的判断是,普通用户并不太关心哪个 frontier model 最强,更关心 onboarding 是否清楚、价格是否合理、产品能否可靠完成工作。因此,营销与产品信息传达的重要性不亚于模型进步。

Sources1

Official Blog

Claude Blog — Building intelligent apps for Apple platforms with Claude in the Foundation Models framework

Anthropic is releasing a Swift package that lets Apple developers call Claude through Apple's Foundation Models framework when a workflow needs multi-step reasoning, code generation, web search, or code execution for data analysis. Apple's framework can handle fast local work such as summarization and extraction, then pass typed Swift values produced through `@Generable` annotations into Claude instead of raw user text. The package returns streamed responses, tool calls, and structured results to the same SwiftUI view, letting one product route each step to the appropriate model. Anthropic says support will be available the following day on iOS 27, iPadOS 27, macOS 27, visionOS 27, and watchOS 27, and requires an Anthropic API key.

Sources1

Anthropic 正在发布一个 Swift package,让 Apple 开发者可以通过 Apple 的 Foundation Models framework 调用 Claude,以处理需要多步 reasoning、代码生成、web search 或通过执行代码完成数据分析的 workflow。Apple 的 framework 可以先处理 summarization、extraction 等快速本地任务,再把通过 `@Generable` annotation 生成的 typed Swift value 传给 Claude,而不是直接传递未经整理的用户文本。这个 package 会把流式响应、tool call 与结构化结果返回同一个 SwiftUI view,使单一产品能够为每一步选择合适的模型。Anthropic 表示,该支持将在次日上线,适用于 iOS 27、iPadOS 27、macOS 27、visionOS 27 和 watchOS 27,并需要 Anthropic API key。

Sources1

Podcast

The MAD Podcast with Matt Turck — How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis

The takeaway: reliable long-horizon agents need explicit behaviors, process-level supervision, and a well-designed information environment, not just successful end-state evals. Basis co-founder Mitch Troyanovsky explains this through accounting agents that can prepare complex tax returns end to end over hours or days. A completed task is not an opaque answer; it includes major decisions, assumptions, and review points, much like a well-structured engineering change. Basis encodes deterministic checks where possible, adds judges for professional judgments, and studies human accounting processes to define what good work looks like. Troyanovsky's warning is direct: even if 100 evals all pass, that does not prove the system will generalize to production if it arrived at the right answers through an unacceptable process.

Sources1

核心结论:可靠的 long-horizon agent 需要明确的 behavior、process-level supervision 与经过设计的信息环境,而不能只依赖最终结果通过 eval。Basis 联合创始人 Mitch Troyanovsky 以 accounting agent 为例解释了这一点:这些 agent 可以连续运行数小时甚至数天,端到端完成复杂的 tax return。任务完成后交付的不应是一个不透明答案,而应包括关键决策、假设与需要审核的位置,就像结构良好的工程改动。Basis 在可能的地方加入 deterministic check,在需要专业判断的地方使用 judge,并研究人类会计流程来定义什么是好工作。Troyanovsky 的警告非常直接:即使 100 个 eval 全部通过,如果系统是通过不可接受的过程得到正确答案,也不能证明它能泛化到 production。

Sources1

Troyanovsky treats context as runtime training data. Because an LLM has a large working memory but no default short-term or long-term memory, long-running agents must leave useful state for their future inference steps, organize that state into an ergonomic ontology, and distinguish canonical company knowledge from stale records. He says the English in an agent system can be more precious than the code because context directly changes runtime performance. Basis is focused on generating better behavioral signals and improving the harness before updating model weights; Troyanovsky expects much of today's context engineering to be absorbed by models in less than five years but not less than two. He does not view a technical trick as the durable moat. The business moat comes from moving quickly, winning market share, and becoming embedded in customer workflows.

Sources1

Troyanovsky 把 context 视为 runtime training data。由于 LLM 拥有很大的 working memory,却默认没有真正的短期或长期记忆,长时间运行的 agent 必须为未来的 inference step 留下有用状态,把这些状态组织成易于 agent 使用的 ontology,并区分公司当前的 canonical knowledge 与已经过时的记录。他认为,agent 系统里的英文有时比 code 更珍贵,因为 context 会直接改变 runtime performance。Basis 目前专注于生成更好的 behavioral signal,并先改进 harness,而不是直接更新模型权重;Troyanovsky 预计,今天大量 context engineering 最终会在五年以内被模型吸收,但不会在两年以内完成。他也不把某个技术诀窍视为持久 moat。真正的 business moat 来自快速行动、赢得市场份额,并深度嵌入客户 workflow。

Sources1

Engineering & Research

Better idea capture and traceable agent writing favor conversation over polished briefs

Meta AI Senior Director Madhu Guru asks teams to record themselves explaining a new idea as they would to a friend, then use AI for basic cleanup while preserving the original structure and flow. His observation is that people are often clearer when speaking; the core idea gets buried when they add context and polish to make a document sound intelligent. Swyx's `ai-devblog` skill applies a related principle to technical writing: it first elicits what the author believes the story is, then works with them to trace what they read and report it faithfully, including visuals. Both workflows use AI to preserve and substantiate a human point of view rather than replace it with generic prose.

Sources12

Meta AI 高级总监 Madhu Guru 要求团队先像向朋友解释一样,把自己对新想法的说明录下来,再让 AI 做基础清理,同时保留原始结构与表达流向。他观察到,人们在说出想法时往往更清楚;一旦开始写文档,为了补充背景、润色并让文字显得聪明,核心观点反而容易被埋没。Swyx 的 `ai-devblog` skill 把类似原则应用到了技术写作:它先引导作者说清楚自己认为真正的故事是什么,再与作者一起追溯阅读材料并忠实报告,同时还能制作视觉内容。这两类 workflow 都让 AI 帮助保留并验证人的观点,而不是用通用文本替代它。

Sources12
Generated through the Follow Builders skill — bilingual daily signal and weekly perspective from the people building AI.