Top Signals
Agent harnesses are becoming portable model interfaces
OpenAI's Thibault Sottiaux said GPT-5.6 Sol can run broadly, including inside the Claude Code harness, and reset usage limits for all paid ChatGPT Work and Codex users. When a user reported an account ban after pairing Anthropic's harness with another model, Sottiaux called that outcome odd and asked whether others had seen the same issue. Together, the posts show coding harnesses separating from their original model providers, while access policies have not fully caught up with that portability.
OpenAI 的 Thibault Sottiaux 表示,GPT-5.6 Sol 可以在多种环境中运行,包括 Claude Code harness,并为所有 ChatGPT Work 和 Codex 付费用户重置了用量限制。当一名用户称自己因在 Anthropic harness 中接入其他模型而被封禁时,Sottiaux 认为这种结果很奇怪,并询问是否还有人遇到同类问题。这两条动态共同表明,coding harness 正在与原始模型提供商解耦,但访问政策尚未完全适应这种可移植性。
Enterprise automation will often be invisible workflow infrastructure
Box CEO Aaron Levie argues that enterprise AI productivity will vary far more than expected because frontier gains require workflows redesigned around agents, a level of complexity most users will not absorb themselves. His prescription is to put agents into existing systems of record and automate the ten highest-leverage processes in the background, rather than waiting for employees to change how they work or encouraging undirected experimentation across the company.
Box CEO Aaron Levie 认为,企业 AI 带来的生产力提升差异会远超预期,因为要释放 frontier 能力,企业必须围绕 agent 重构 workflow,而这种复杂度不可能由大多数用户自行承担。他的建议是,把 agent 接入现有 system of record,在后台自动化最具杠杆效应的十个流程,而不是等待员工改变工作方式,或在全公司推动缺乏方向的自由实验。
Production agents need retrieval discipline, tools, and narrow scope
AI educator Peter Yang says the largest bottleneck in building a strong agent is often not the model. Teams weaken agents by burying them in excessive context, withholding tools that let them retrieve what they need, and attempting broad coverage before mastering a few core use cases. His account of Linear's production-agent process points toward a product discipline built around selective context and explicit tool access rather than ever-larger prompts.
AI 教育者 Peter Yang 表示,打造优秀 agent 的最大瓶颈往往不是模型。团队会因为塞入过量 context、没有提供自主检索所需信息的工具,以及在做好少数核心场景之前就追求大而全,而削弱 agent 的表现。他对 Linear 生产级 agent 流程的概括,指向一种以精选 context 和明确 tool access 为核心的产品纪律,而不是不断扩大 prompt。
AI may consume the apprenticeship pipeline it still depends on
Builder Zara Zhang highlights the “Tragedy of the Cognitive Commons”: checking AI output requires expertise, but expertise is built through years of junior-level work, precisely the work AI is removing first. Each company can rationally eliminate entry-level tasks while the profession collectively loses its mechanism for training people who can catch future model errors. The unresolved governance question is not only who reviews AI today, but who will still be qualified to review it in fifteen years.
Builder Zara Zhang 强调了“认知公地悲剧”:检查 AI 输出需要专业能力,而专业能力来自多年初级工作积累,恰恰是 AI 最先取代的那部分工作。每家公司都可能理性地削减 entry-level 任务,但整个行业会因此失去培养未来模型纠错者的机制。尚未解决的治理问题不只是今天由谁审核 AI,更是十五年后还有谁具备审核资格。
Agent-era cloud controls are becoming machine-readable
Vercel CEO Guillermo Rauch outlined a cost and abuse-control stack spanning soft and hard spending caps, anomaly alerts, recursion protection for Functions, always-on DDoS mitigation, and billing APIs that agents can query directly. He says the controls depend on streaming infrastructure that analyzes a very large volume of data in real time to detect anomalies and threats and dispatch alerts. Making usage and cost data accessible to agents turns financial guardrails into part of the runtime control plane, not just a human-facing billing dashboard.
Vercel CEO Guillermo Rauch 介绍了一套成本与滥用控制体系,包括软硬支出上限、异常告警、Functions 递归保护、持续开启的 DDoS 防护,以及可由 agent 直接查询的账单 API。他表示,这些能力依赖 streaming infrastructure,对大量数据进行实时分析,以发现异常与威胁并发送告警。让 agent 能直接读取用量和成本数据,意味着财务 guardrail 正在成为 runtime control plane 的一部分,而不再只是面向人的账单 dashboard。
The social license for AI infrastructure cannot be assumed
FirstMark investor Matt Turck argues that resistance to data centers is not simply irrational NIMBYism. Nearby communities may distrust coastal technology companies, see construction jobs as temporary, fear being left with unfinished facilities if the AI boom reverses, find little daily value in AI, and still face real electricity-generation impacts even when water systems are closed-loop. His warning is that dismissing these concerns as foreign influence or ignorance will deepen the technology industry's political problem.
FirstMark 投资人 Matt Turck 认为,社区对 data center 的抵制不能简单归结为不理性的 NIMBYism。当地居民可能不信任来自沿海地区的科技公司,认为建设岗位只是暂时的,担心 AI 热潮退去后留下未完工设施,也没有在日常生活中感受到 AI 的明显价值;即便用水采用闭环系统,发电带来的影响仍然真实存在。他警告说,如果把这些担忧斥为外国影响或无知,只会进一步加剧科技行业的政治困境。
Podcast
Chasing Trillion-Dollar Companies, Founder Ambition, Token Budgets, and Regulatory Capture with Sarah & Elad
The Takeaway: Sarah and Elad argue that AI's extraordinary upside does not erase the need for disciplined arithmetic: founders and investors must distinguish enormous markets from businesses capable of reaching $50 billion to $100 billion in revenue, then repeatedly reassess whether a company is capturing value as capabilities rise and costs fall.
Gil describes the recent cluster of trillion-dollar outcomes as a punctuated technology wave, not proof that every promising robotics, materials, or AI application company can reach the same scale within three to five years. Guo pushes back that investors still underestimate outcome-based business models, but both see a troubling drop in founder ambition as teams retreat into niches they assume frontier labs will ignore. For most companies, they recommend a pre-scheduled, unemotional board discussion every six to twelve months about whether to sell. The calculation should include future dilution, realistic revenue multiples, financing structure, and the founder's finite time: “your most productive years of your life are on the line right now.”
The same scarcity logic is moving inside AI labs. Compute, not researcher salaries, is becoming the binding constraint, so labs increasingly allocate token budgets toward the small number of people and projects expected to produce outsized results. Gil calls the emerging metric “return on invested tokens” and expects enterprises to move from broad AI experimentation toward measuring spend, using more open-source systems, and concentrating compute on core products or major margin improvements. Physical compute constraints may also reinforce an oligopoly by limiting how quickly any one lab can pull ahead.
Their policy concern is regulatory capture disguised as one-sided safety analysis. They argue society must compare risk with benefit, especially where faster AI progress could improve healthcare, education, productivity, transportation, and elder care. The position is not that safeguards are unnecessary, but that excessive restrictions can protect incumbents and suppress beneficial deployment, repeating patterns they identify in energy and biotechnology.
核心结论: Sarah 和 Elad 认为,AI 的巨大上行空间并不能取代严谨的商业计算。创始人与投资人必须区分“市场规模巨大”和“单家公司能达到 500 亿至 1000 亿美元收入”,并随着能力提升、成本下降,持续重新判断公司是否真正捕获了价值。
Gil 把近期集中出现的万亿美元级公司视为一次技术浪潮中的间断式爆发,而不是每一家有前景的机器人、材料或 AI 应用公司都能在三到五年内达到同等规模的证据。Guo 则反驳说,投资人仍然低估了按结果收费的商业模式;但两人都注意到一个令人担忧的趋势:一些团队因为认定 frontier lab 不会进入某些细分领域,反而降低了创业雄心。对于大多数公司,他们建议每六到十二个月预先安排一次不带情绪的董事会讨论,认真评估是否出售。计算时应纳入未来稀释、现实收入倍数、融资结构,以及创始人有限的时间,因为“你人生中最有生产力的岁月,此刻正押在这件事上”。
同样的稀缺性逻辑也正在进入 AI lab。真正的约束越来越不是研究员薪资,而是 compute,因此 lab 会把更多 token budget 分配给预期能产生超额成果的少数人才与项目。Gil 将正在形成的指标称为“return on invested tokens”,并预计企业会从广泛试用 AI,转向衡量支出、更多采用 open-source 系统,并把 compute 集中到核心产品或显著改善利润率的项目。物理 compute 约束还可能强化寡头格局,因为它限制了任何一家 lab 拉开差距的速度。
他们对政策的担忧是,被包装成单边安全分析的 regulatory capture。社会必须同时比较风险与收益,尤其是在更快的 AI 进展可能改善医疗、教育、生产力、交通和养老服务的领域。他们并非主张取消 safeguard,而是认为过度限制会保护 incumbent、压制有益部署,并重演他们在能源和生物科技领域看到的模式。