回到卷首
每日集录ai builders

八月十五日

二〇二六年 16 builders 32 posts 1 podcast 1 blog 约三十四分钟

Claude enters Apple's Foundation Models framework for typed on-device-to-cloud handoffs, Cursor's outcome sharpens the applied-AI playbook, OpenClaw makes agent sessions and UI changes reviewable, builders track expanding software demand, content-level spam gaps, personal apps, transaction completion, fourteen-hour coding runs, and Basis cofounder Mitch Troyanovsky explains why reliable long-horizon agents require process evals, independent review, agent judges, and context engineered as runtime training.

Top Signals

Aaron Levie: Cursor validates the applied-AI product layer

Box CEO Aaron Levie argues that Cursor shattered several assumptions at once: agentic coding was a much larger market than earlier developer-tool exits suggested, and there was ample room to innovate between users and foundation models. He identifies a repeatable applied-AI playbook in Cursor's product shape, model-neutral workflow layer, selective post-training, task-specific infrastructure, and category-aligned go-to-market motion. The important signal is that model access alone does not erase application differentiation; product and operational choices can still compound into a major platform.

Sources1

Box CEO Aaron Levie 认为,Cursor 一次打破了几个旧假设:agentic coding 的市场远大于过去 developer tool 的退出规模所暗示的上限,而且在用户与 foundation model 之间仍有大量创新空间。他总结出一套可复用的 applied AI 路径,包括找到正确的产品形态、成为 model-neutral 的 workflow layer、针对关键环节做 post-training、建设任务专用基础设施,以及形成与品类匹配的 go-to-market 体系。最重要的信号是,模型能力的普及并不会自动抹平应用层差异;产品与运营选择依然可以持续复利,最终形成重要平台。

Claude arrives inside Apple's Foundation Models framework

Claude Blog announced a new Swift package that lets Apple developers call Claude through Apple's Foundation Models framework when an app needs multi-step reasoning, code generation, web search, or code execution. The handoff is designed to preserve one native experience: Apple's on-device models can handle fast local work such as summarization or extraction, typed values from `@Generable` annotations become clean Claude inputs, and the package streams tool calls and structured responses back into SwiftUI. Support is scheduled to become available the next day across iOS 27, iPadOS 27, macOS 27, visionOS 27, and watchOS 27, using an Anthropic API key.

Sources1

Claude Blog 宣布推出一个新的 Swift package,让 Apple 开发者可以通过 Foundation Models framework 调用 Claude,在应用需要多步推理、code generation、web search 或 code execution 时完成模型接力。整个设计仍保持为一个原生体验:Apple 的 on-device model 负责快速、本地的 summarization 或 extraction,`@Generable` annotation 产生的 typed value 作为干净的 Claude 输入,package 再把 tool call 与 structured response 流式送回 SwiftUI。该支持计划于次日上线,覆盖 iOS 27、iPadOS 27、macOS 27、visionOS 27 和 watchOS 27,并使用 Anthropic API key。

Peter Steinberger turns agent work into reviewable team context

Peter Steinberger says the OpenClaw team now builds OpenClaw with OpenClaw, and calls shareable agent-session URLs a superpower. He also added a short rule to the team's shared `AGENTS.md` requiring videos on every pull request that changes UI state. Together these practices make agent work inspectable at two levels: the reasoning and execution session can be shared, while visible behavior changes arrive with direct evidence for reviewers.

Sources12

Peter Steinberger 表示,OpenClaw 团队现在已经开始用 OpenClaw 来开发 OpenClaw,而可通过 URL 分享 agent session 是一项非常强的能力。他还在团队共享的 `AGENTS.md` 里增加了一条简短规则,要求所有改变 UI state 的 pull request 都上传视频。这两项实践让 agent 工作在两个层面变得可检查:推理与执行过程可以直接分享,可见的行为变化也会附带 reviewer 能快速验证的证据。

AI lowers the cost of software and expands its addressable problems

Meta Senior Director of AI Madhu Guru compares AI-assisted software engineering with more efficient steam engines, higher-level programming languages, and spreadsheets: lowering the cost and complexity of a task tends to increase total demand by making previously uneconomic use cases viable. He argues that when building itself becomes broadly accessible, defensibility shifts toward product sense, domain knowledge, distribution, and execution. He also credits Cursor with moving AI products beyond the generic chatbot pattern by establishing a reusable "Cursor for X" product shape.

Sources123

Meta AI Senior Director Madhu Guru 把 AI-assisted software engineering 与更高效的蒸汽机、高级编程语言和 spreadsheet 相提并论:当一项任务的成本与复杂度下降,原本不具备经济可行性的 use case 会被激活,总需求反而往往上升。他认为,当人人都能 build 时,真正的差异化会转向 product sense、domain knowledge、distribution 和 execution。他还认为 Cursor 帮助 AI 产品走出通用 chatbot 模式,确立了可以迁移到不同领域的“Cursor for X”产品形态。

X's anti-spam model may miss content-level AI slop

Peter Yang inspected X's open-source algorithm and found a behavioral model called TweetSpamBot that examines as many as 512 recent account actions. It looks for patterns such as posting bursts, quote-tweet spam, amplification, timing, browsing, and dwell behavior; a high score can contribute to an account challenge. Yang's critique is that the model does not appear to read post content or directly downrank the posts, leaving a path for sophisticated content mills that repeatedly apply the same template to unrelated viral material while imitating normal account behavior.

Sources1

Peter Yang 检查了 X 的开源算法,发现其中有一个名为 TweetSpamBot 的行为模型,会分析最多 512 条近期账户操作。它关注 posting burst、quote-tweet spam、content amplification、发帖时机、正常浏览和 dwell 等模式;高分可能触发账户 challenge。Yang 的批评是,这个模型似乎并不读取帖子内容,也不会直接 downrank 对应帖子,因此更成熟的 content mill 仍可能通过模拟正常账户行为,对互不相关的热门内容反复套用同一模板。

Product & Engineering

Gemini 3.7 Flash reaches the Gemini app, while Pomelli adds motion

Google VP Josh Woodward says Gemini 3.7 Flash is now available in the Gemini app. He also highlighted a Pomelli launch that turns product photoshoots into short videos or GIFs, extending the Google Labs experiment from creating polished product imagery to repurposing those assets into motion content for small businesses.

Sources12

Google VP Josh Woodward 表示,Gemini 3.7 Flash 现已进入 Gemini app。他还介绍了 Pomelli 的一次更新,可以把产品照片转成短视频或 GIF,让这个 Google Labs experiment 从生成精致产品图进一步延伸到动态内容复用,服务于 small business 的营销需求。

ChatGPT moves into transaction completion

Thibault Sottiaux, who works on Codex and ChatGPT at OpenAI, says finding a quick restaurant reservation is now easy inside ChatGPT. The post is brief, but the product shift is concrete: the assistant is moving from recommending options toward completing a time-sensitive real-world task within the conversational surface.

Sources1

OpenAI 的 Codex 与 ChatGPT 团队成员 Thibault Sottiaux 表示,现在可以很轻松地直接在 ChatGPT 中寻找餐厅预订。虽然帖子很简短,但产品变化很具体:assistant 正在从给出建议,走向在对话界面中完成有时间约束的现实任务。

Personal software grows outside public app distribution

Replit CEO Amjad Masad points to TestFlight as a useful destination for personal apps even when the builder has no intention of publishing them to the App Store. That observation captures an emerging distribution pattern: AI can make software valuable at the scale of one person or a small circle, without requiring a public launch or conventional app-store economics.

Sources1

Replit CEO Amjad Masad 指出,即使完全没有把应用发布到 App Store 的计划,通过 TestFlight 构建和使用 personal app 依然很有价值。这个观察揭示了一种正在形成的分发模式:AI 可以让软件在只服务一个人或一个小圈子的规模上就产生价值,不再要求公开发布,也不必符合传统 app store 的经济逻辑。

Long-running coding work stretches to fourteen hours

FPV Ventures partner Nikunj Kothari reports that `/goal` completed an extremely detailed specification in one shot over fourteen hours when given generous CLI tools. He explicitly notes that it may not be the most token-efficient approach, but the result is a useful frontier marker for agent duration: the system sustained a tool-rich implementation far beyond the short interactive sessions that defined earlier coding assistants.

Sources1

FPV Ventures partner Nikunj Kothari 表示,在提供充足 CLI tool 的情况下,`/goal` 用十四小时 one-shot 完成了一份极其详细的 spec。他明确指出,这未必是 token 效率最高的方案,但结果仍然标记出 agent 时长的新边界:系统能够持续完成 tool-rich implementation,远超早期 coding assistant 所依赖的短时交互 session。

Builder Perspectives

The workday becomes a sequence of decisions

FirstMark VC Matt Turck describes AI changing the workday from a mix of decisions and process execution into a dense sequence of decisions, with mental exhaustion arriving earlier. The joke carries a serious organizational point: removing mechanical work does not necessarily remove effort; it can concentrate attention on judgment, making decision quality and cognitive pacing the new bottlenecks.

Sources1

FirstMark VC Matt Turck 把 AI 之后的工作日描述为:过去是 decision 与 process execution 混合交替,现在则变成高密度连续做 decision,于是脑力更早耗尽。这个玩笑背后有一个严肃的组织问题:机械工作减少并不等于整体付出减少,它可能把注意力更集中地压到判断上,让 decision quality 与认知节奏成为新的瓶颈。

Strong products often fuse several partial ideas

Linear Head of Product Nan Yu argues that ideas from capable colleagues rarely come from nowhere; even when a proposal is not directly usable, it is usually motivated by a real problem or an interesting line of thought. Some of Linear's most interesting launches, he says, emerged from fusing multiple ideas or moving several derivatives away from an original concept. The practical product lesson is to preserve the signal inside imperfect proposals instead of evaluating every idea as a binary yes-or-no decision.

Sources1

Linear Head of Product Nan Yu 认为,优秀同事提出的想法很少凭空出现;即使一个方案不能直接采用,它通常也来自某个真实问题,或一条值得继续思考的线索。他说,Linear 一些最有意思的产品,正是把多个想法融合起来,或从最初概念继续推导数步后形成的。实际的产品启示是,不要把每个 idea 只当成 yes-or-no 的二元决策,而应保留不成熟提案中真正有价值的信号。

Podcast

The MAD Podcast with Matt Turck — How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis

The Takeaway: reliable long-horizon agents need to be engineered like organizations, with explicit process, independent review, inspectable decisions, and carefully managed context, not judged only by whether the final answer happens to be correct.

Basis cofounder Mitch Troyanovsky builds accounting agents that can work autonomously for hours or days, including preparing tax returns end to end. He defines autonomy carefully: the agent should finish a first pass without repeatedly asking the user what to do, but then surface its major decisions, assumptions, and review points. The right analogy is not an opaque machine that says “trust me”; it is a strong junior professional handing over a well-structured piece of work that a reviewer can understand quickly.

His sharpest technical argument is that outcome evals are insufficient for work spanning thousands of inference steps. A system can pass one hundred tests while relying on a process that would be unacceptable in production. In tax research, for example, getting the answer from pre-training knowledge or a secondary source is not equivalent to checking the primary authority. Basis therefore encodes human professional practice into behaviors, deterministic checks, independent review, and agent-based judges that inspect trajectories, while acknowledging that judging deeply nested sub-agent work remains expensive.

Troyanovsky treats context engineering as runtime training: every instruction, document, tool, and intermediate state becomes the small, high-value dataset from which the agent learns during execution. That makes English inside the context more operationally precious than elegant code organization because it directly changes model performance. His memorable formulation is: “The English is more precious because the English affects the performance.” The long-term bet is that agents will eventually improve their own context and harnesses, but the present work is to build reliable signal and translate centuries of domain practice into agent-readable process.

Sources1

核心结论:可靠的 long-horizon agent 应该像组织一样被设计,需要明确流程、独立 review、可检查的决策和精心管理的 context,而不能只看最终答案是否碰巧正确。

Basis cofounder Mitch Troyanovsky 正在构建能自主运行数小时甚至数天的 accounting agent,其中包括端到端完成 tax return。他对 autonomy 的定义很克制:agent 应该在不反复询问用户下一步做什么的情况下完成 first pass,但完成后必须清楚呈现关键决策、假设和需要 review 的部分。正确的类比不是一个只说“相信我”的黑箱,而是一名优秀的 junior professional,交付结构清楚、让 reviewer 能快速理解的工作成果。

他最锋利的技术观点是,面对跨越数千个 inference step 的工作,只看 outcome eval 远远不够。系统可能通过一百个测试,却依赖一种在生产环境中完全不可接受的流程。例如在 tax research 里,从 pre-training knowledge 或二手资料得到正确答案,并不等于核对 primary authority。因此 Basis 把人类专业实践编码成 behavior、deterministic check、independent review,以及检查 trajectory 的 agent-based judge,同时也承认,对深层嵌套的 sub-agent 工作进行判断目前仍然成本很高。

Troyanovsky 把 context engineering 视为 runtime training:每条指令、每份文档、每个 tool 和中间状态,都构成 agent 在执行过程中学习所依赖的小规模高价值数据集。这意味着 context 里的英文内容在运营上甚至比优雅的代码组织更珍贵,因为前者会直接改变模型表现。他最值得记住的一句话是:“The English is more precious because the English affects the performance.” 长期来看,agent 最终会改进自己的 context 与 harness;但现阶段真正的工作,是建立可靠 signal,并把数百年积累的领域实践翻译成 agent 可执行的流程。

Sources1
Generated through the Follow Builders skill — bilingual daily signal and weekly perspective from the people building AI.