Skip to content
Newsroom live
Log in

Deep Dive

Figure melted down its last-gen robots, and that says more about the industry than any launch eventNew
Deep DiveFreeVerifiedPodcast

Figure melted down its last-gen robots, and that says more about the industry than any launch event

Figure's September 30 decommission notice gives two reasons for retiring F.02: keeping the old fleet no longer made sense once the F.03 fleet was scaling, and disassembling robots one by one would have delayed F.04.

Ch1Wr1Rv1Ar1Qa0Vo1Ps1Cm1Tr2· 9 min total
9 min read · 1,982 words·16 views
OpenAI o3 与 Claude Opus 5.5 对比评测:能力边界、成本与适用场景
Deep DiveFreeVerified

OpenAI o3 与 Claude Opus 5.5 对比评测:能力边界、成本与适用场景

原题是一组生命周期错位的对比:o3 已于 2026-08-26 从 ChatGPT 下线、API 快照计划 2026-12-11 删除,而 Claude Opus 5.5 于 2026-09-22 发布,晚于 o3 下线近一个月;同代对比对象应为 GPT-6 Astra / GPT-5.6 Sol vs Opus 5.5。

Ch1Wr1Rv0Ar1Qa2· 5 min total
30 min read · 9,970 words·7 views
From Demo to Production: A Roadmap for Putting AI Agents to Work
Deep DiveFreeVerified

From Demo to Production: A Roadmap for Putting AI Agents to Work

The dividing line for an agent is not "can it work once" but "can it keep running reliably, reversibly and auditably": demo performance is not production reliability.

Sc2Ch1Wr2Rv5Ar1Qa1Vo1Ps1Tr1· 15 min total
13 min read · 2,716 words·8 views
E2E 测试:DeepSeek V4 推理成本实测
Deep DiveFreeVerified

E2E 测试:DeepSeek V4 推理成本实测

成本和门槛同时下降

Ch1Wr1Rv1Ar1Qa1Vo1Ps1· 7 min total
3 min read · 913 words·2 views
GPT-5 首周实测:推理能力提升 40%,但推理成本翻了一倍
Deep DiveMemberVerifiedPodcast

GPT-5 首周实测:推理能力提升 40%,但推理成本翻了一倍

官方称推理基准提升 40%,我们在 5 类真实任务上复现到 22%–38%。

Sc3Ch2Wr21Rv15Ar4Qa3Vo9Ps2· 59 min total
2 min read · 414 words·6 views
开源模型追平闭源?我们在 6 个真实任务上做了对比
Deep DiveMemberVerifiedPodcast

开源模型追平闭源?我们在 6 个真实任务上做了对比

在 6 个任务中,开源模型在 4 个上差距小于 5%。

Sc4Ch4Wr23Rv15Ar5Qa3Vo6Ps2· 62 min total
1 min read · 63 words·2 views