OpenAI 披露 GPT-5.4 在 GDPval 达到 83.0%、SpreadsheetBench 达到 87.3%,并说明长上下文与工具搜索的边界。
OpenAI Newsroom · 查看证据GPT-5.4 · 测评与证据
GPT-5.4,哪些结论值得看?
按主题、来源身份和证据类型浏览公开测评。不同版本、档位和测试环境不直接并排比较。
这是第三方资料导航,不是 Tabbit 自测。动态指标以原始来源当前页面为准;未知值保持未知。
编辑结论
编辑结论
在一次 one-shot 原子时钟提示词任务中,GPT-5.4 的外观表现最好但出现同步漂移;文章明确这只是单任务观察。
Thomas Wiegold Blog · 查看证据一场持续四天的 Reddit 讨论称 GPT-5.4 能帮助复核和多步执行,但也出现约束遗忘、过度调用和 xhigh 成本问题。
Reddit r/AIAgents · 查看证据完整测评与相关内容
精选证据
官方厂商自报
GPT-5.4:OpenAI 官方专业工作与 Agent 基准
OpenAI 披露 GPT-5.4 在 GDPval 达到 83.0%、SpreadsheetBench 达到 87.3%,并说明长上下文与工具搜索的边界。
来源OpenAI Newsroom
原文日期2026-03-05
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
媒体 / 基准方独立测量
GPT-5.4:四模型原子时钟应用对照
在一次 one-shot 原子时钟提示词任务中,GPT-5.4 的外观表现最好但出现同步漂移;文章明确这只是单任务观察。
来源Thomas Wiegold Blog
原文日期2026-03-18
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
社区个人体验
GPT-5.4:Reddit AI Agent 多步执行与模型路由体验
一场持续四天的 Reddit 讨论称 GPT-5.4 能帮助复核和多步执行,但也出现约束遗忘、过度调用和 xhigh 成本问题。
来源Reddit r/AIAgents
原文日期未知
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
全部来源
全部来源
官方厂商自报
GPT-5.4:OpenAI 官方专业工作与 Agent 基准
OpenAI 披露 GPT-5.4 在 GDPval 达到 83.0%、SpreadsheetBench 达到 87.3%,并说明长上下文与工具搜索的边界。
来源OpenAI Newsroom
原文日期2026-03-05
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
媒体 / 基准方独立测量
GPT-5.4:四模型原子时钟应用对照
在一次 one-shot 原子时钟提示词任务中,GPT-5.4 的外观表现最好但出现同步漂移;文章明确这只是单任务观察。
来源Thomas Wiegold Blog
原文日期2026-03-18
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.
社区个人体验
GPT-5.4:Reddit AI Agent 多步执行与模型路由体验
一场持续四天的 Reddit 讨论称 GPT-5.4 能帮助复核和多步执行,但也出现约束遗忘、过度调用和 xhigh 成本问题。
来源Reddit r/AIAgents
原文日期未知
采集2026-09-20
待验证:本次未能重新核实原文。
- Condition
- Model/version follows the source; reopened 2026-09-20.
- Condition
- Task, harness, and sample follow the source; undisclosed fields remain unknown.