回到卷首
每日集录ai builders

八月二十六日

二〇二六年 17 builders 32 posts 1 podcast 约二十六分钟

Claude unifies memory across Chat and Cowork with explicit controls for saved and sensitive topics, Vercel ships secure connectivity and lightweight code execution for agents, builders surface a macOS keychain risk and turn evolving workflows into eval roadmaps, applied AI differentiates through enterprise integration, and Parallel founder Parag Agrawal lays out agent-native search and a value-based economic model for the web.

Top Signals

Claude unifies memory across Chat and Cowork, with user-visible controls

Anthropic's Claude account said Claude now carries one memory across Chat and Claude Cowork, so a Cowork task can begin with context already established in chat. Saved memory appears as a list of topics in Settings where users can read, edit, or delete each item; it also updates during conversations, and users can explicitly say "remember this." Memory is on by default for Free, Pro, and Max plans, while topics Anthropic describes as sensitive, including health and religious beliefs, remain excluded unless the user enables them. Cat Wu said the memory system is now unified across Chat and Cowork because of user feedback, and Boris Cherny described the update as simpler and more powerful.

Anthropic 的 Claude 账号表示,Claude 现在会在 Chat 与 Claude Cowork 之间共享同一套 memory,因此 Cowork 接到任务时,可以直接使用此前在 chat 中已经建立的上下文。保存的 memory 会以 topic 列表呈现在 Settings 中,用户可以逐项读取、编辑或删除;它也会随着对话自动更新,用户还可以明确说出 "remember this" 来保存特定内容。Free、Pro 和 Max 方案默认开启 memory,但 Anthropic 认定的敏感主题,包括健康和宗教信仰,除非用户主动启用,否则不会被写入。Cat Wu 表示,这次统一 Chat 与 Cowork 的 memory 更新来自用户反馈;Boris Cherny 则把这次变化概括为更简单也更强大。

Sources12345

Vercel ships two lighter-weight building blocks for connected agents

Vercel CEO Guillermo Rauch announced that Vercel Connect is generally available, framing secure connectivity to services and data as the hardest problem in agent development. A developer can create a connection such as Notion from the CLI and receive an MCP client that queries on behalf of the authenticated user. He also introduced Run SDK for evaluating dynamically generated Code Mode programs inside a lightweight QuickJS secure context, positioning it as a faster and more cost-efficient option when an agent's code does not need a full sandbox.

Vercel CEO Guillermo Rauch 宣布 Vercel Connect 已正式 GA,并把安全连接 services 与 data 称为 agent 开发中最难的问题。开发者可以通过 CLI 创建 Notion 等连接,然后获得一个 MCP client,代表已经完成身份验证的用户发起查询。他还发布了 Run SDK,用于在轻量级 QuickJS secure context 中 eval Code Mode 动态生成的程序;当 agent 写出的代码不需要完整 sandbox 时,这是一种更快、成本更低的选择。

Sources12

A Codex macOS capability triggers a keychain lockout warning

Swyx warned users to avoid the Codex capability he called “locked use” for now. He said it relies on unstable macOS features and had completely locked him out of his macOS keychain twice in one week; he also pointed to an Apple developer-forum acknowledgment that the underlying issue is a known bug. The report is a concrete reminder that agent permissions and operating-system integration can fail below the model or harness layer.

Swyx 提醒用户,现阶段应避免使用他所称的 Codex “locked use” capability。他表示,该功能依赖不稳定的 macOS features,并在一周内两次让他完全无法访问 macOS keychain;他还提到,Apple developer forum 已确认底层问题属于 known bug。这份报告再次具体说明,agent permissions 与操作系统集成可能在 model 或 harness 之外的更底层发生故障。

Sources1

Applied AI's moat is the workflow around the model, Aaron Levie argues

Box CEO Aaron Levie argued that a large gap remains between AI models and the workflows of an enterprise, leaving substantial room for applied AI companies. He said the value lies in understanding domain context and change management, routing across models, connecting critical business systems, fitting agents into the right user experience, and building domain-specific evals. His core claim is that customers buy resolved problems and achieved outcomes, not raw tokens, models, or agents.

Box CEO Aaron Levie 认为,AI models 与企业实际 workflows 之间仍存在巨大鸿沟,这为 applied AI 公司留下了可观空间。他认为,真正的价值来自理解 domain context 与 change management、在不同 models 之间进行 routing、连接关键业务系统、让 agents 以合适的用户体验进入 workflow,以及建立特定领域的 evals。他的核心判断是,客户购买的是问题得到解决和结果真正达成,而不是原始 tokens、models 或 agents。

Sources1

Engineering & Research

Madhu Guru says evals need a product roadmap, not a frozen test set

Meta Senior Director of AI Madhu Guru argued that evals fail when teams treat them as static while user behavior advances from short summaries to multi-document synthesis and eventually proactive monitoring. His financial-research example moves from one five-page earnings report to five reports, then 15 filings, transcripts, and research reports, and finally portfolio monitoring. Each step changes what must be tested: short to long context, single-turn to multi-turn work, passage citations to document-and-line citations, simple QA to complex synthesis, and reactive chat to proactive agents. He recommends mapping the dimensions of expected usage, mining production traces, prioritizing the next stage's P0 evals, and iterating on observed failure modes before customers get there.

Meta Senior Director of AI Madhu Guru 认为,如果团队把 evals 当成静态资产,而用户行为已经从短篇摘要发展到多文档综合,最终走向主动监控,那么 evals 就会失效。他以金融研究为例:需求会从总结一份五页财报,发展到分析五份财报,再到综合 15 份 filings、transcripts 和 research reports,最后变成持续监控投资组合。每个阶段都要求不同的测试重点:从 short context 到 long context、single-turn 到 multi-turn、passage citations 到 document-and-line citations、simple QA 到 complex synthesis,以及 reactive chat 到 proactive agents。他建议先描绘预期使用方式的演进维度,分析 production traces,为下一阶段建立 P0 evals,再根据实际 failure modes 持续迭代,尽量走在用户需求前面。

Sources1

An open-source skill turns fragmented cancer-care information into one working brief

Peter Yang, who publishes practical AI tutorials and interviews, open-sourced `/fuck-cancer`, a skill for patients and caregivers navigating diagnosis and treatment. It builds a living brief from supplied documents and context, organizing patient and care-team details, no more than three next actions, confirmed facts versus open questions, plain-English definitions, and an update-and-decision log. When research is needed, Yang said it uses sources including the National Cancer Institute and the ClinicalTrials.gov API. He uses it with ChatGPT/Codex and Claude Code, with output saved locally as Markdown or shared through an updated Google Doc so a family can maintain one source of truth.

发布实用 AI tutorials 与 interviews 的 Peter Yang 开源了 `/fuck-cancer`,这是一个帮助患者及照护者应对诊断和治疗流程的 skill。它会根据用户提供的 documents 与 context 持续维护一份 brief,其中包括患者与 care team 信息、不超过三项 next actions、已确认事实与待澄清问题、通俗易懂的术语解释,以及更新和决策日志。需要研究时,Yang 表示它会使用 National Cancer Institute、ClinicalTrials.gov API 等来源。他会在 ChatGPT/Codex 和 Claude Code 中使用这项 skill,结果既可以保存为本地 Markdown,也可以更新到共享 Google Doc,让家人围绕同一份 source of truth 协作。

Sources1

Product & Workflow

Google Labs makes vibe coding multiplayer with Putty

Google Labs introduced Play with Putty, an experiment for collaboratively building tools and websites in real time. The product turns the usually solitary vibe-coding loop into a shared session; access begins through a waitlist and is currently limited to people aged 18 or older in the United States.

Google Labs 发布了 Play with Putty,这是一项让多人实时协作构建 tools 和 websites 的实验。它把通常由个人完成的 vibe coding loop 变成共享 session;目前需要加入 waitlist,而且仅面向美国 18 岁及以上用户。

Sources1

Podcast

Training Data: Parallel’s Parag Agrawal: Building a New Web for AI Agents

The Takeaway: Former Twitter CEO and Parallel founder Parag Agrawal is rebuilding web search around agents as the customer, with ranking, economics, and interfaces optimized for machine work rather than human clicks. Parallel began with the bet that agents would use the web 1,000 times more than humans. Agrawal's sharpest formulation is that “human click data is a bug”: an agent doing search should be trained and evaluated with agent feedback because its queries are longer, more explicit, and aimed at retrieving the highest-signal excerpts rather than pages designed for people to click through.

核心结论: 前 Twitter CEO、Parallel founder Parag Agrawal 正在把 agents 视为新的客户,围绕机器执行任务的方式重建 web search,包括 ranking、商业模式和交互接口,而不再以人类点击为中心。Parallel 创立时的核心押注是,agents 使用 web 的规模将达到人类的 1,000 倍。Agrawal 最尖锐的表述是“human click data is a bug”:负责 search 的 agent 应该用 agent feedback 训练和评估,因为它发出的 queries 更长、更明确,目标是拿到 signal 最高的 excerpts,而不是打开为人类点击而设计的页面。

Parallel's initial wedge was not an instant search engine but a search agent that could crawl after a query arrived, spending a minute or sometimes ten minutes on research. That made outsourced human web work the first competitor: insurance underwriting and claims, sales enrichment, and financial data collection. These workloads generated empirical evals and let Parallel expand its index incrementally. The current system rewrites a query across multiple indexes, retrieves from tens or hundreds of billions of URLs, and applies progressively larger ranking stages until it returns roughly the thousand tokens with the most signal. Agrawal said Parallel search generally lets an agent use under half as many tokens while becoming more accurate and faster.

Parallel 最初的切入点并不是即时 search engine,而是可以在收到 query 后开始 crawl 的 search agent,一次研究可持续一分钟,有时甚至十分钟。这样一来,它最先竞争的对象是外包给人类完成的 web 工作,包括保险 underwriting 与 claims、sales enrichment,以及金融数据收集。这些工作提供了真实的 evals,也让 Parallel 能够逐步扩大 index。当前系统会面向多个 indexes 重写 query,从数百亿乃至上千亿 URLs 中进行 retrieval,再通过逐级增大的 ranking stages,最终返回 signal 最高的大约一千个 tokens。Agrawal 表示,使用 Parallel search 通常可以让 agent 消耗不到一半的 tokens,同时获得更高准确率和更快速度。

Agrawal also sees agent traffic breaking the web's ad and subscription economics because publishers cannot monetize or even interpret an anonymous machine visit the way they can a human one. Parallel's proposed alternative uses estimated Shapley values to price a source by the incremental quality it adds, with differentiated payment for unique content and for content used in higher-value work. His longer-term vision moves the web from pull to push: instead of an agent searching only after a request, changes on the web become a feed that wakes agents when something actionable happens.

Agrawal 还认为,agent traffic 正在打破 web 原有的广告与订阅经济,因为 publishers 无法像理解和变现人类访问那样,处理一次匿名机器访问。Parallel 提出的替代方案使用估算的 Shapley values,按照某个 source 带来的增量质量定价,并针对独特内容,以及被用于更高价值工作的内容,实施差异化支付。他更长期的愿景是让 web 从 pull 转向 push:agent 不再只在收到请求后搜索,而是由 web 上的变化形成 feed,在出现可执行事件时主动唤醒 agents。

Sources1
Generated through the Follow Builders skill — bilingual daily signal and weekly perspective from the people building AI.