X / TWITTER
Anthropic's Cat Wu (Claude Code + Cowork) shared one of her favorite real-world Claude Code workflows: recruiting. She describes the role and the backgrounds she's after, then asks Claude Code to kick off a dynamic workflow that sources 100 candidates — each with LinkedIn, Twitter, blog, and podcast links plus a one-line pitch — generate an artifact, and email it to her. Then she locks her laptop, heads out for the day, and reviews the finished list on the go. It's a concrete template for using an agent as an always-on research assistant that hands you a decision-ready deliverable instead of a chat log.
Anthropic 的 Cat Wu(Claude Code + Cowork)分享了她最喜欢的一个 Claude Code 真实工作流:招人。她先把职位要求和想找的候选人背景讲清楚,然后让 Claude Code 启动一个 dynamic workflow 去搜出 100 位候选人,每人都带上 LinkedIn、Twitter、博客和播客链接,外加一句话简介,再生成一个 artifact 并邮件发给她。之后她就锁上电脑出门,等 Claude Code 跑完,在路上直接过这份名单。这给出了一个很具体的模板:把 agent 当成一个全天候的调研助手,最后交给你的是一份可以直接拍板的成品,而不是一堆聊天记录。
Linear head of product Nan Yu pushed back on a popular flex, arguing that "bragging about running 10 Claude Code tabs is just theater." His deeper point: the "real-time strategy game" model of managing agents — a human frantically micro-managing many parallel agents by hand — is a dead end. He notes that AI which is "ancient by current standards" already beats more than 99% of human RTS players and out-micros them to an extreme degree, so the human-at-the-controls, babysitting-many-tabs paradigm is exactly where the leverage won't be.
Linear 产品负责人 Nan Yu 泼了一盆冷水,他认为"炫耀自己同时开着 10 个 Claude Code 标签页,纯粹是在演戏"。他更深一层的观点是:把管理 agent 当成一场"即时战略游戏"(real-time strategy game)来玩,也就是一个人手忙脚乱地同时微操一堆并行 agent,这条路是走不通的。他指出,哪怕是"按今天标准已经很古老"的 AI,也早就能击败超过 99% 的人类 RTS 玩家,并在微操上把人类甩开一大截。所以"人坐在操控台前、盯着一堆标签页当保姆"这套范式,恰恰是最不可能产生真正杠杆的地方。
Y Combinator CEO Garry Tan argued that AI has removed the real bottleneck on wealth creation: "the real constraint on human wealth was never resources. It was good ideas for how to serve one another, and the leverage to act on them. We just deleted the leverage constraint for everybody. Now it's only the ideas." Writing from Osaka, he extended the thought to Japan, which he says already "ran the experiment on the ceiling" — thirty years of zero growth, yet the best trains, service, and craft on earth. His lesson: "when you can't compete on more, you compete on better," and AI now lets builders do better and more at the same time.
Y Combinator CEO Garry Tan 认为 AI 已经拆掉了创造财富的真正瓶颈:"人类财富的真正约束从来不是资源,而是'如何相互服务'的好点子,以及把它付诸行动的杠杆。我们刚刚替每个人删掉了杠杆这个约束,现在剩下的就只有点子了。"人在大阪的他把这个想法延伸到日本,说日本其实早就"跑完了触及天花板的那场实验",三十年零增长,却造出了地球上最好的铁路、服务和手艺。他的结论是:"当你没法在'更多'上竞争时,就在'更好'上竞争",而 AI 让 builder 现在可以同时做到更好和更多。
FPV Ventures partner Nikunj Kothari vented about the ritual of fundraising Zoom calls, saying he's "still waiting on that one founder who gates a VC chat until they've played with the product and come with at least two pieces of feedback." Talking through the same parroted deck for 30 minutes, he argues, is far less productive than a joint product brainstorm — or simply getting to know each other. His half-joking punchline: both founders and VCs could just upload their personal "prompts" to Claude and let it run the repetitive pitch conversation digitally.
FPV Ventures 合伙人 Nikunj Kothari 吐槽了融资 Zoom 会议这套仪式,他说自己"还在等这么一位创始人:非要等到对方玩过产品、并带着至少两条反馈,才肯开始跟 VC 聊"。他认为,对着同一份背得滚瓜烂熟的 deck 讲上 30 分钟,远不如一起做一场产品头脑风暴来得高效,或者干脆就好好互相认识一下。他半开玩笑地收尾:不如创始人和 VC 双方都把各自的"prompt"上传给 Claude,让它把这段重复的 pitch 对话在线上跑完算了。
OpenAI CEO Sam Altman marveled that his older kid put two words together for the first time, saying he's "approximately as amazed by this cognitive feat as I am by GPT-5.6 discovering new math." Wrapped in a personal aside is a rare public signal from Altman that OpenAI's latest model is producing genuinely novel mathematical results — a claim that lines up with reports of frontier models cracking previously open problems.
OpenAI CEO Sam Altman 感叹自己的大孩子第一次把两个词连在一起说出口,还说他"对这个认知成就的惊叹,大概和我对 GPT-5.6 发现新数学的惊叹差不多"。这句看似私人的随口感慨里,藏着 Altman 难得的一次公开信号:OpenAI 最新的模型正在产出真正意义上全新的数学结果,这也和 frontier 模型攻克此前未解难题的一些说法对得上。
PODCASTS
No Priors — Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
The Takeaway: In the reasoning era, a model's intelligence is no longer a fixed number — it's a function of how much compute you're willing to spend at inference time, and almost everyone is still measuring it wrong.
Noam Brown, the OpenAI research scientist who pioneered inference-time scaling, argues the industry's beloved "benchmark grid" (one score per model, per benchmark) has become actively misleading. When GPT-5.5 launched, the grid showed only a few points of gain over 5.4 and sparked skepticism. The real story: 5.5 is far more efficient with its thinking, and once you control for test-time compute, the jump is substantial. His fix is blunt — stop publishing a single number and instead plot performance against a budget axis of tokens, dollars, or time.
The deeper problem is that modern models don't plateau. Scaffold 5.5 well and it can think productively "for weeks even," which means nobody actually knows the ceiling of models that are already released. He points to OpenAI's disproof of the Erdős unit-distance conjecture, done at a budget that was "dirt cheap" — and notes that people later coaxed the same result out of public 5.5 with the right scaffolding, roughly $100k of compute that no one had bothered to spend. His memorable line on the pace: put a hard problem down, "come back two months later, and then it's a thousand times cheaper."
The uncomfortable mirror is safety. If capability keeps climbing with budget, so do dangerous capabilities, yet frameworks built in the ChatGPT era "don't really account for the amount of test time compute." The only way to truly know what an agent does after running for a year may be to run it for a year — impossible under a two-month release cycle.
On recursive self-improvement, Brown is a skeptic of overnight takeoff: because peak intelligence demands enormous compute, "time itself becomes a bottleneck." Models still lack research taste — they optimize brilliantly but can't yet invent genuinely novel algorithms — so for now they amplify researchers rather than replace them.
The Takeaway:在推理(reasoning)时代,一个模型的智能不再是一个固定的数字,而是你愿意在推理时投入多少算力的函数,而几乎所有人至今仍在用错误的方式衡量它。
Noam Brown,这位开创了 inference-time scaling 的 OpenAI 研究科学家认为,业界钟爱的"benchmark 网格"(每个模型在每个 benchmark 上给一个分数)如今正在实实在在地误导人。GPT-5.5 发布时,网格上显示它只比 5.4 高出几个百分点,于是招来了质疑。但真实情况是:5.5 的思考效率高得多,一旦把 test-time compute 拉齐来比较,这个提升其实相当可观。他给出的解法很直接:别再只公布一个数字,而要把性能画成一条随预算(token、美元或时间)变化的曲线。
更深层的问题是,现代模型根本不会"见顶"。只要 scaffold 得当,5.5 甚至能"连续高效思考好几周",这意味着没有人真正知道那些已经发布的模型的能力上限在哪里。他举了 OpenAI 推翻 Erdős 单位距离猜想的例子,说这件事花的预算"便宜到不像话",而且后来有人用合适的 scaffold,也从公开版 5.5 里把同样的结果"引导"了出来,大约 10 万美元的算力,只是此前没人愿意花而已。他有一句关于节奏的名言:把一个难题先放一放,"两个月后再回来,成本就已经便宜了一千倍"。
而这面镜子令人不安的另一面是安全。如果能力会随预算不断攀升,那么危险能力同样会,可 ChatGPT 时代搭建的那些框架"根本没有把 test-time compute 的量考虑进去"。要真正弄清一个 agent 跑满一年之后能做什么,唯一可靠的办法可能就是真的让它跑满一年,而这在两个月一次的发布周期下是不可能的。
关于递归自我改进(recursive self-improvement),Brown 并不相信会有一夜之间的"起飞":因为要达到最高智能就得依赖海量算力,于是"时间本身成了瓶颈"。模型目前仍然缺乏研究品味(research taste),它们能把算法优化得极其漂亮,却还发明不出真正新颖的算法,所以眼下它们是在放大研究者的能力,而不是取代研究者。