官方把 V3.2 定位为日常平衡型、可在 thinking 与 non-thinking 中调用工具的 Agent 模型,把 V3.2-Speciale 定位为最高推理/竞赛型但发布时不支持工具,二者不能混作一个模型。
模型:DeepSeek-V3.2 与 DeepSeek-V3.2-Speciale;V3.2 可用 App/Web/API,Speciale API-only。
能力:推理、工具使用、Agent 数据合成、竞赛数学/编程;官方发布页没有逐项公开完整 harness。
上下文/成本:发布说明强调 V3.2 的 balanced inference vs length;没有在该页列出完整 token/价格表。
官方称 V3.2 继承 V3.2-Exp 使用方式;V3.2-Speciale 当时通过临时 endpoint 提供,到 2025-12-15 结束,same pricing、no tool calls。V3.2 的 thinking/tool-use 详情指向 Thinking Mode 文档。
V3.2:官方称 balanced inference vs length,作为日常 driver,性能达到 GPT-5 level(官方定位语句)。
V3.2-Speciale:官方称 maxed-out reasoning,接近 Gemini-3.0-Pro,并在 IMO、CMO、ICPC World Finals、IOI 2025 达到 gold-level results。
Agent 训练:覆盖 1,800+ environments 和 85k+ complex instructions;V3.2 首次把 thinking 直接整合进 tool-use,并支持 thinking/non-thinking 工具模式。
如果工作流需要 API 工具调用、日常编码和研究 Agent,选择 V3.2 并固定 thinking 状态;如果只比较数学/极限推理,可单独研究 Speciale,但不能把 Speciale 的竞赛成绩当作 V3.2 Agent 能力。
发布页是厂商定位与选择性结果,没有完整原始输入、失败样本、成本和方差。
“GPT-5 level”“rivals Gemini-3.0-Pro”是官方描述,不是同 harness 的独立证明。
Speciale 临时 endpoint 已过期(按发布页时间),不能据此设计当前生产接入。
Agent 数据合成规模不是 benchmark score。
锁定当前可用模型 ID;不要使用过期 Speciale endpoint。
对 V3.2 的 thinking/non-thinking 分别跑相同任务,记录工具调用、reasoning_content、token、延迟和最终正确性。
对数学/编程任务另建竞赛 split,明确是否允许工具和额外 token。
报告 V3.2 Agent 与 Speciale reasoning 两条结果线,不合并成一个总分。
DeepSeek V3.2