X / Twitter
Aaron Levie (Box CEO)
Box CEO Aaron Levie argued that multi-model agentic systems are clearly the future, pointing to new research from Cursor showing that pairing a frontier model as planner and orchestrator with a cheaper workhorse model can cut total token costs on a project by 15X. He quotes the key insight: "Few moments in a large task genuinely require frontier intelligence, such as the original decomposition, the design decisions, and certain trade-offs. Once a frontier planner has collapsed the ambiguity into a detailed, explicit instruction, less expensive models simply have to follow it." He sees this becoming the core design pattern of complex agents — and the template for how the applied layer differentiates: companies that know a domain well and can route across model tiers in coding, finance, legal, healthcare, and life sciences will win workloads that would otherwise be too expensive for customers to deploy.
Box CEO Aaron Levie 认为多模型 agent 系统显然是未来方向,他引用了 Cursor 的最新研究:用前沿模型做规划和编排、用更便宜的模型干活,可以把一个项目的总 token 成本降低 15 倍。他引用了其中的关键洞察:"大型任务中真正需要前沿智能的时刻很少,比如最初的任务拆解、设计决策和某些权衡。一旦前沿规划者把模糊性收敛成详细、明确的指令,便宜的模型照着执行就行了。"他认为这正在成为复杂 agent 的核心设计模式,也是 AI 应用层建立差异化的模板:真正懂行业、又能在多个模型档位之间做路由的公司,将在编程、金融、法律、医疗、生命科学等关键领域拿下那些原本因为太贵而无法部署的工作负载。
Guillermo Rauch (Vercel CEO)
Vercel CEO Guillermo Rauch distilled what he calls the big lesson from AI: everything is code. A slide deck is code. Design is code. That cool promo video? Code. Excel automation? Code. "The universe? Probably made of code too."
Vercel CEO Guillermo Rauch 总结了他所说的 AI 带来的最大启示:一切皆代码。幻灯片是代码,设计是代码,那个很酷的宣传视频是代码,Excel 自动化也是代码。"宇宙?大概也是代码写的。"
Swyx (Latent Space / Cognition)
Swyx flagged a notable trajectory-comparison writeup buried in the RLM paper from researchers Alex Zhang and Omar Khattab, calling out an open secret of frontier model training: even without training on the test set, labs can effectively cheat by training on test lookalikes, letting them goalseek almost any benchmark number. And since open-weight releases almost never ship the datasets or RL environments that would reveal it, there is built-in plausible deniability. The authors explore applying standard NLP distance metrics to hidden trajectories — no ultimate solution yet, but their preliminary analysis supports the finding that RLMs can genuinely generalize to unseen tasks that share latent structure with training.
Swyx 特别推荐了藏在 RLM 论文里的一段轨迹对比分析,作者是研究者 Alex Zhang 和 Omar Khattab。他点破了前沿模型训练的一个公开秘密:即使不直接在测试集上训练,也可以通过训练"测试集高仿题"来变相作弊,想刷出什么 benchmark 分数几乎都能实现。而开源权重发布时 99% 的情况都不会附带能揭穿这件事的数据集或 RL 环境,所以天然存在推诿空间。论文作者尝试用标准 NLP 距离度量来分析隐藏轨迹,虽然还没有终极解法,但初步探索支持了一个结论:RLM 确实能泛化到与训练数据共享潜在结构的未见任务上。
Peter Yang (Creator, AI educator)
Creator Peter Yang made a pointed prediction: banning Chinese models will be the same self-own as banning Chinese EVs. He also shared a practical agent pattern from Anthropic's Thariq Shihipar: use one agent to do the work and a separate agent to review it against a rubric, because of what they call self-preferential bias — "When a model prefers its own output, it's going to be more lenient at verifying it." For subjective outputs like video shorts with no deterministic answer, a separate verification agent reading a rubric gives far more honest feedback.
创作者 Peter Yang 给出了一个尖锐的预测:封禁中国模型将会和封禁中国电动车一样,是一次自伤行为。他还分享了来自 Anthropic 的 Thariq Shihipar 的一个实用 agent 模式:让一个 agent 干活,再用另一个独立的 agent 对照评分标准做审查,原因是所谓的自我偏好偏差,"当模型偏爱自己的输出时,它在验证自己的输出时就会更宽松。"对于视频短片这类没有确定性答案的主观产出,让一个独立的验证 agent 读评分标准再给反馈,得到的意见要诚实得多。
Zara Zhang (Builder)
Builder Zara Zhang laid out how she would structure hiring today: Round 1 is in-person with no AI allowed, testing domain expertise and on-the-fly thinking; Round 2 is a project that is impossible to complete without AI, where the candidate is assessed not just on the result but on their chat transcript with agents. She also observed that there are now basically two kinds of companies — those built before coding agents, scrambling to retrofit, and those founded after, which are different from day one: teams under ten people because they genuinely don't need more, work organized by projects instead of departments, each person closing their own loop, and almost no internal meetings.
Builder Zara Zhang 给出了她如果今天招人会怎么设计面试流程:第一轮线下面试,禁止使用 AI,现场考察领域专业度和临场思考;第二轮是一个不用 AI 根本不可能完成的项目,考察的不只是最终结果,还包括候选人与 agent 的完整对话记录。她还观察到,现在基本上只有两种公司:一种是在 coding agent 出现之前创立的,正在手忙脚乱地改造自己;另一种是之后创立的,从第一天起就不一样:团队不到十个人,因为真的不需要更多;工作按项目而非部门组织;每个人自己闭环;几乎没有内部会议。
Nikunj Kothari (FPV Ventures Partner)
FPV Ventures partner Nikunj Kothari warned that many founders of the last 18 months are about to learn that "no moats in AI" does not mean scale and capital become your moat instead. History is full of companies that were structurally and financially in great positions — Webvan, Groupon, MySpace, Yahoo, AltaVista, Blockbuster, Nokia, the zombiecorns of 2021 — and each eventually got beat by a company with a much better unique insight, or collapsed under the weight of its own scale. The needle founders have to thread: find a unique insight worth a 10+ year journey, while being prudent enough to not let capital and scale become a substitute for it.
FPV Ventures 合伙人 Nikunj Kothari 提醒,过去 18 个月入场的许多创始人很快会意识到,"AI 时代没有护城河"并不意味着规模和资本就自动成为你的护城河。历史上有太多结构和资金层面都处于绝佳位置的公司:Webvan、Groupon、MySpace、Yahoo、AltaVista、Blockbuster、Nokia,还有 2021 年那批僵尸独角兽,它们最终要么被拥有更好独特洞察的公司击败,要么被自身规模的重量压垮。创始人必须穿过的针眼是:找到一个值得投入十年以上的独特洞察,同时保持足够的克制,不让资本和规模成为洞察的替代品。
Madhu Guru (Meta Sr Director of AI)
Meta Senior Director of AI Madhu Guru argued that the road to AGI is paved with economically valuable tasks — which is exactly why enterprise AI is one of the most important frontiers, because that's where many of those tasks live. He also noted the irony that four years past the peak of web3 tokenomics debates, the tokenomics debate that actually matters turned out to be open versus closed weights, inference costs, and model routing. And with models this capable, he declared it's literally the greatest time ever to have product sense.
Meta AI 高级总监 Madhu Guru 认为,通往 AGI 的道路是由具有经济价值的任务铺成的,这正是企业级 AI 是最重要前沿之一的原因,因为大量这样的任务就在企业场景里。他还指出了一个讽刺之处:web3 tokenomics 大辩论过去四年后,真正重要的 tokenomics 辩论原来是开源权重与闭源权重之争、推理成本和模型路由。他还感慨,在模型如此强大的当下,现在简直是拥有产品直觉的最好时代。
Amjad Masad (Replit CEO)
Replit CEO Amjad Masad spotlighted what he asked may be the first physical product shipped by a coding agent.
Replit CEO Amjad Masad 转发并提问:这是不是第一个由 coding agent 交付的实体产品?
Podcasts
No Priors — Travel Through the Lens of AI with with Booking.com CEO Glenn Fogel
The Takeaway: The CEO moving $186 billion a year in travel says even his own scale is no protection — "There is no such thing as a moat. There is no such thing as somewhere you're gonna be protected against innovation."
Glenn Fogel runs Booking Holdings, and his career is a masterclass in surviving bubbles: he joined Priceline in early 2000, days before the Nasdaq peaked, watched the stock fall to a dollar a share, and stayed twenty-seven years as it climbed roughly a thousandfold to near $6,000. He sees clear parallels between 1999 and the AI boom: enormous real value being created, and a great deal of capital about to be destroyed, with no way to guess the ratio in advance.
His sharpest point is how badly outsiders read travel. When OpenAI dropped its ChatGPT checkout feature, Booking's stock jumped 8% — but he thinks both the original panic and the relief were overreactions from people who don't understand the business: travel is heavily regulated worldwide and getting more so, and servicing hotel partners takes thousands of people, not just inventory in a database. "If you think you're just gonna come in and do this business and knock away these very big players, I'd say you should really understand what the business is before you decide to commit your capital."
Meanwhile he's building the disruption himself. He used Penny, Priceline's agentic assistant, to plan a complex multi-city Europe family trip — two cabin classes, one kid flying back to a different city, frequent flyer miles versus cash per leg — and it worked through the trade-offs like a concierge. Penny adoption has doubled monthly with higher conversion and lower cancellations, yet he's deliberately not pushing it hard because token economics remain unsolved: cost per booked trip, which model for which purpose, and whether the lifetime value justifies it. AI customer service has already cut cost per contact while satisfaction rose.
On jobs, his worry isn't displacement itself but the speed mismatch between jobs destroyed and jobs created: booking.com once employed humans across 40 languages for translation, and machine translation erased all of it. His answer is relentless upskilling — because fear-driven rejection of AI would hobble the West while China charges ahead.
核心要点:这位每年经手 1860 亿美元旅行交易的 CEO 说,即使是他自己的规模也不是保护伞,"世上根本不存在护城河。不存在任何一个能让你免受创新冲击的地方。"
Glenn Fogel 执掌 Booking Holdings,他的职业生涯堪称穿越泡沫的教科书:2000 年初,就在 Nasdaq 见顶前几天,他加入了 Priceline,眼看着股价跌到 1 美元,然后坚守了二十七年,见证它上涨约一千倍、逼近 6000 美元。他认为 1999 年与这轮 AI 热潮有清晰的相似之处:真实价值在大量创造,同时也有大量资本即将灰飞烟灭,而成败比例事先无从预测。
他最犀利的观点是外行对旅行业的误读有多严重。OpenAI 砍掉 ChatGPT 的 checkout 功能时,Booking 股价应声上涨 8%,但他认为最初的恐慌和之后的如释重负都是不懂这门生意的人的过度反应:旅行业在全球范围内受到严格监管且监管还在加码,服务酒店合作伙伴需要成千上万的人力,而不只是把库存塞进数据库。"如果你以为自己进场就能干掉这些巨头,我会说,在决定投入资本之前,你真的应该先搞懂这门生意是什么。"
与此同时,他自己正在亲手打造这场颠覆。他用 Priceline 的 agentic 助手 Penny 规划了一次复杂的欧洲多城市家庭旅行:两种舱位、一个孩子要飞回另一座城市、每段航程用里程还是付现金,Penny 像一位私人旅行管家一样逐项权衡。Penny 的使用量连续数月每月翻倍,转化率更高、取消率更低,但他刻意没有全力推广,因为 token 经济学还没算清楚:每促成一单旅行的成本是多少、什么场景该用什么模型、用户终身价值是否划算。AI 客服则已经降低了单次联系成本,同时满意度还在上升。
关于就业,他担心的不是岗位消失本身,而是岗位消失与新岗位创造之间的速度错配:booking.com 曾雇佣覆盖 40 种语言的人工翻译团队,机器翻译让这一切彻底消失。他的答案是不停歇的技能升级,因为出于恐惧而拒绝 AI 只会让西方自缚手脚,而中国不会停下来。