Deep Dive
Deep DiveFreeVerifiedPodcast
Figure melted down its last-gen robots, and that says more about the industry than any launch event
Figure's September 30 decommission notice gives two reasons for retiring F.02: keeping the old fleet no longer made sense once the F.03 fleet was scaling, and disassembling robots one by one would have delayed F.04.
Ch1Wr1Rv1Ar1Qa0Vo1Ps1Cm1Tr2· 9 min total
9 min read · 1,982 words·16 views
Deep DiveFreeVerified
OpenAI o3 与 Claude Opus 5.5 对比评测:能力边界、成本与适用场景
原题是一组生命周期错位的对比:o3 已于 2026-08-26 从 ChatGPT 下线、API 快照计划 2026-12-11 删除,而 Claude Opus 5.5 于 2026-09-22 发布,晚于 o3 下线近一个月;同代对比对象应为 GPT-6 Astra / GPT-5.6 Sol vs Opus 5.5。
Ch1Wr1Rv0Ar1Qa2· 5 min total
30 min read · 9,970 words·7 views
Deep DiveFreeVerified
From Demo to Production: A Roadmap for Putting AI Agents to Work
The dividing line for an agent is not "can it work once" but "can it keep running reliably, reversibly and auditably": demo performance is not production reliability.
Sc2Ch1Wr2Rv5Ar1Qa1Vo1Ps1Tr1· 15 min total
13 min read · 2,716 words·8 viewsDeep DiveFreeVerified
E2E 测试:DeepSeek V4 推理成本实测
成本和门槛同时下降
Ch1Wr1Rv1Ar1Qa1Vo1Ps1· 7 min total
3 min read · 913 words·2 viewsDeep DiveMemberVerifiedPodcast
GPT-5 首周实测:推理能力提升 40%,但推理成本翻了一倍
官方称推理基准提升 40%,我们在 5 类真实任务上复现到 22%–38%。
Sc3Ch2Wr21Rv15Ar4Qa3Vo9Ps2· 59 min total
2 min read · 414 words·6 viewsDeep DiveMemberVerifiedPodcast
开源模型追平闭源?我们在 6 个真实任务上做了对比
在 6 个任务中,开源模型在 4 个上差距小于 5%。
Sc4Ch4Wr23Rv15Ar5Qa3Vo6Ps2· 62 min total
1 min read · 63 words·2 views