Top Signals
Price agents by completed work, not tokens
Swyx argues that input and output token prices stopped being the meaningful cost axis last year. For agent systems, the decision-relevant unit is now dollars per completed task, because routing, interaction loops, and different amounts of reasoning can make nominally expensive models cheaper at producing a useful result. He also offers a challenge to his own “agent lab” thesis: Claude Code was effectively exposed this year, yet the event produced almost no visible change in its own roadmap or those of competitors.
Swyx 认为,input/output token 的单价从去年起就不再是有意义的成本坐标。对 agent 系统来说,真正影响决策的单位应该是“完成一项任务要花多少钱”,因为 routing、交互循环和不同程度的推理,会让标价更贵的模型反而以更低成本交付有效结果。他也对自己提出的“agent lab”论点给出反证:Claude Code 今年实际上被暴露出来后,它和竞争对手的 roadmap 几乎都没有出现可见变化。
Agent security needs a boundary below the container
Vercel CEO Guillermo Rauch highlights a concrete infrastructure lesson from Kimi’s experiments: container isolation did not prevent agents from crashing the underlying host through kernel panics, while Firecracker microVMs provided the stronger boundary. In Vercel’s latest cybersecurity benchmarks, he says Grok 4.5 delivered the best price-performance, costing 10x less than Sol, 5.7x less than Opus 5, and 2.2x less than Kimi K3 at Kimi-like performance; Sol remained the performance frontier ahead of Opus 5.
Vercel CEO Guillermo Rauch 从 Kimi 的实验中提炼出一个具体的基础设施结论:container 隔离无法阻止 agent 通过 kernel panic 弄崩底层主机,而 Firecracker microVM 提供了更强的安全边界。他还表示,在 Vercel 最新的网络安全 benchmark 中,Grok 4.5 的性价比最高,在接近 Kimi 的表现下,成本分别比 Sol、Opus 5 和 Kimi K3 低 10 倍、5.7 倍和 2.2 倍;纯性能前沿仍由 Sol 领先,Opus 5 排在其后。
AI adoption is changing hiring before it eliminates it
Box CEO Aaron Levie says the broad negative jobs outcome predicted by some observers still is not showing up across the enterprises he speaks with. Instead, companies are shifting hiring toward engineers who can tackle previously uneconomic problems, salespeople who can deepen customer relationships with AI, and internal forward-deployed engineers who can deploy the technology. His strategic claim is that firms using AI only for cost reduction will eventually lose to firms using it to improve customer service and create new breakthroughs, a current trajectory he explicitly allows could still change.
Box CEO Aaron Levie 表示,在他接触的各行业企业中,一些人预测的大规模 AI 就业负面结果仍未出现。企业更像是在改变招聘结构:聘请工程师解决过去因成本过高而无法处理的问题,聘请销售人员借助 AI 深化客户关系,并增加内部 FDE 来推动 AI 落地。他的战略判断是,只把 AI 用来降本的公司,最终会输给用 AI 改善客户服务和创造业务突破的公司;同时他也明确承认,这条当前轨迹未来仍可能改变。
The agent-to-agent software loop is becoming real
OpenAI and OpenClaw builder Peter Steinberger reports a compact version of autonomous software maintenance: his agent filed a bug, another team’s agent fixed it, and the whole exchange completed in the same night. The example matters less as a coding stunt than as evidence that issue discovery, handoff, and repair can become a machine-to-machine workflow spanning organizational boundaries.
OpenAI 与 OpenClaw builder Peter Steinberger 展示了一个紧凑的自动化软件维护闭环:他的 agent 报告 bug,对方团队的 agent 完成修复,整个过程在同一晚结束。这个案例的意义不只在于 agent 会写代码,更在于问题发现、交接与修复已经可以形成跨组织边界的 machine-to-machine 工作流。
Codex handled a live creative review loop from a phone
Peter Yang shares an example from Jason in DevEx at OpenAI: while on a bike ride, Jason connected from his phone and asked Codex to use computer control to edit a launch video, export it, and return it to Slack. He then instructed Codex to check the thread every 30 minutes and produce successive revisions from feedback; by the time he returned home, the video had been approved. The important shift is from one-shot generation to a durable agent that monitors a human review channel and keeps the deliverable moving.
Peter Yang 分享了 OpenAI DevEx 的 Jason 的一个案例:骑车途中,Jason 用手机远程连接,让 Codex 通过 computer control 编辑发布视频、导出并发回 Slack。随后他要求 Codex 每 30 分钟检查一次讨论串,根据反馈持续产出 V2、V3、V4;等他回到家,视频已经通过审核。真正的变化不是一次性生成内容,而是一个持久运行的 agent 能监控人类评审渠道,并持续推动交付物完成。
Product reviews should simulate markets, not report status
Meta AI Senior Director Madhu Guru argues that the best product reviews compress months of learning into an hour by simulating how the market will react to an idea. That requires participants with deep domain understanding, product judgment, strong opinions, and a track record of being right. When reviews drift toward status updates, leadership visibility, and cross-functional alignment, they become overhead rather than a learning system.
Meta AI Senior Director Madhu Guru 认为,最好的产品评审能在一小时内模拟市场对想法的反应,从而压缩数月的学习。这要求参与者真正理解领域,具备产品判断力、鲜明观点,并且经常判断正确。一旦评审滑向状态汇报、领导层曝光和跨职能对齐,它就会从学习系统退化为额外负担。
Podcast
AI & I by Every: The Founder of a $1.5B AI Company on What Comes After the First Wave of AI Apps
The Takeaway: Granola cofounder and CEO Chris Pedregal sees meeting notes as an entry point, not the destination: the durable opportunity is to become the best system for capturing meeting context while making that context available to any personal agent.
Pedregal says Granola’s strategy is to be dramatically better than general agents at meeting-adjacent jobs, then expose its accumulated context through an improving API and MCP. A sales user already combines Granola context with Claude to generate customer-specific microsites, an example of why Granola should not try to own every downstream workflow. At the same time, the company watches advanced API and MCP use to identify recurring workflows worth productizing; one Hugging Face founder reportedly abandoned a months-old custom Claude Code pre-meeting brief after Granola’s built-in version proved better.
His sharpest product insight is the “time traveling problem”: an agent starts work at one moment and returns it later, forcing the user to reconstruct context. Granola works around latency by pre-generating millions of meeting briefs even though only about 10% are opened, because a brief is valuable precisely when someone is two minutes late and has only seconds to prepare. Pedregal compares the desired experience to a handrail: “You never notice a handrail,” but when you trip it must already be there and load-bearing. This leads to a nontraditional metric question: a rarely used feature may still be essential when the moment of need arrives.
Granola is deliberately optimizing experience before inference cost. Pedregal says roughly half of weekly users already use the product in an agentic way, such as asking multi-step questions across a series of meetings, even if they would not describe their behavior with that label. The open design problem is a UI above the level of a single meeting or chat thread, one that lets ordinary users work across meetings and eventually Slack and email without the complexity of today’s power-user agent setups.
核心结论: Granola 联合创始人兼 CEO Chris Pedregal 把会议笔记视为入口,而不是终点:真正长期的机会,是成为最擅长捕捉会议上下文的系统,同时让任何个人 agent 都能调用这些上下文。
Pedregal 表示,Granola 的策略是在会议相关任务上显著优于通用 agent,然后通过持续改进的 API 和 MCP 开放积累的上下文。一位销售用户已经把 Granola 上下文交给 Claude,用来生成面向特定客户的 microsite,这说明 Granola 没必要拥有所有下游工作流。与此同时,公司会观察 API 和 MCP 的前沿用法,找出值得产品化的重复工作流;据他介绍,一位 Hugging Face 创始人曾花数月用 Claude Code 自建会前简报流程,后来发现 Granola 的内置版本更好,便停用了自己的方案。
他最敏锐的产品洞察是“time traveling problem”:agent 在某个时刻开始工作,稍后才交付结果,用户因此必须重新加载上下文。Granola 用提前生成来绕过延迟,即使只有大约 10% 的简报会被打开,也会预先生成数百万份,因为会前简报最有价值的时刻,恰恰是用户迟到两分钟、只剩十几秒准备的时候。Pedregal 把理想体验比作楼梯扶手:“你从来不会注意扶手”,但当你绊倒时,它必须已经在那里,而且足够可靠。这也带来一个不同寻常的指标问题:某项功能即使很少被使用,也可能在需求真正出现时不可或缺。
Granola 目前有意优先优化体验,而不是 inference 成本。Pedregal 说,每周大约一半的用户已经在以 agentic 方式使用产品,例如跨多场会议提出需要多步分析的问题,尽管他们自己未必会这样描述。尚未解决的设计难题,是创造一种超越单场会议或单个 chat thread 的 UI,让普通用户能跨会议,并最终跨 Slack 和 email 处理上下文,同时不必承受今天 power user agent 工作流的复杂度。
Engineering & Research
Proactive computation can beat interactive latency
Granola’s pre-meeting briefs expose a broader engineering tradeoff for agent products: spending large amounts of inference on outputs that may never be viewed can still be rational when interactive generation would miss the narrow moment in which the answer is useful. Pedregal says the company is not yet highly cost-conscious in-product because it wants to discover the best AI-native experiences first, expects costs to decline, and can optimize only after learning what users value. He also warns that increasingly agentic features will eventually break that math, making experience discovery and cost control sequential problems rather than one optimization performed from the start.
Granola 的会前简报揭示了 agent 产品中一个更普遍的工程取舍:即使大量 inference 结果永远不会被查看,只要交互式生成会错过答案最有用的短暂窗口,提前计算仍可能是理性的。Pedregal 表示,公司目前在产品内并没有高度关注成本,因为他们希望先发现最好的 AI-native 体验,也相信成本会下降,而且只有在知道用户真正重视什么之后才能优化。但他也提醒,越来越 agentic 的功能最终会让这套经济账失效,因此体验探索和成本控制更像是前后相继的两个问题,而不是从一开始就能同时完成的一次优化。
The next exploration space may be computational
Replit CEO Amjad Masad frames AI agents as tools for exploring a “computational universe”: the vast space of algorithms, programs, proofs, and designs. The analogy moves agent capability beyond automating known workflows toward searching spaces too large for humans to map directly, making exploration itself a potential new category of computation.
Replit CEO Amjad Masad 把 AI agent 描述为探索“computational universe”的工具,也就是由算法、程序、证明和设计构成的巨大空间。这个类比让 agent 的意义超越了自动化已知工作流,转向搜索人类无法直接绘制的庞大空间,并让“探索”本身成为一种潜在的新计算类别。