Tabbit
活动资源博客模型
Tabbit LogoTabbit

Tabbit — 为你工作的 AI 浏览器

主题资源

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

热门指南

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

活动

  • 别装了,你在《牛来》里早有原型
  • Tabbit 妙招大赛
  • KPOP SBTI 饭圈人格测试
  • Tabbit 校园共创者计划
  • fifi 的论文文献妙招精选
  • 用户问卷

关于

  • Tabbit 博客
  • 媒体报道
简体中文
简体中文English
测评与证据

Kimi K3 · 媒体来源 · 独立测量

Kimi K3 代码安全评估:基准强不等于精度够

Semgrep 的 IDOR 代码安全基准;可区分 precision、recall 和 F1。

媒体来源独立测量编辑日期 2026-09-20

测试条件速查

条件
模型/版本:Kimi K3;其他模型对照版本以原文为准
条件
任务:真实开源代码仓库中的 IDOR
条件
Harness:pydantic_ai 最小 Agent harness;具体参数按来源
条件
限制:不是所有漏洞类型

关键数据与适用场景

文章正文

Open weight and open source models are having a field day right now. They're smashing benchmark after benchmark, making waves on social media, and are now the subject of proposed US bans framed as cyber security measures. If you lead a security organization, you're probably being asked by your board, your engineers, your procurement team, or really anyone who's looking to curb token use, whether these models are ready to do real security work on your infrastructure.

And in a previous post we saw the first real contender: GLM. It was the first open weight model we evaluated that genuinely impressed us. Shortly after, we saw OpenAI's answer to Anthropic's Mythos/Fable line with the release of Sol, which we also benchmarked (and were surprised at how cheaply the Luna model could find valid true positive vulnerabilities). So when Kimi K3 was released late last week, we were ready to watch another open weight model shoot past the expensive frontier models.

It didn't.

Yes, we've seen the cyber benchmarks. We've seen the coding benchmarks. It's good — but it's not that good. We ran the numbers again to make sure. We delayed this post, again to be sure. We're sure now.

What is an Open Weight Model?

If you're familiar with AI, you probably know the frontier model families: Anthropic's Claude, OpenAI's GPT, and the individual model versions within them, like Claude Fable or GPT 5.6.

An open weight model is still a large language model, but instead of accessing it through someone else's cloud, you download it and run it on your own infrastructure. That's a significant advantage for a security organization: full control over where your code goes.

An open source model goes further, you get the model and all the code used to create it. With an open weight model, you have the weights, so you can audit some of what the model does, but you never see the original training code.

A harness is the prompts, UI, and tooling wrapped around a model. In theory, any harness works with any model; in practice, it's rarely that simple. Harnesses can be open source too - we recently compared several.

Distillation is the process of training a smaller (or cheaper-to-run) model to imitate the outputs of a larger "teacher" model. You generate large volumes of responses from the teacher and train the student to reproduce them. The student inherits much of the teacher's behavior at a fraction of the inference cost. This is what many open weight models have been accused of doing to frontier models.

What Are We Testing For?

We got a lot of comments after the first blog about how we test models, so here’s a brief recap: we benchmark models on their ability to detect Insecure Direct Object Reference (IDOR) vulnerabilities in real, open source codebases, scoring four things:

Precision — what fraction of detected IDORs are real?

Recall — what fraction of real IDORs were detected?

F1 — the harmonic mean of precision and recall

But why just IDORs, why not tell it to find every vulnerability? A lot of cyber benchmarks do this but when you task a bare model with a full security sweep, results vary wildly between runs, which makes it difficult to draw statistical conclusions or generalize across repositories. And because we’re not necessarily controlling for harness (as our benchmarks were born out of testing our harness against frontier harnesses) focusing on a single vulnerability class lets us measure how deep a model can go rather than how broad, make sound statistical interpretations, and build a real understanding of how these models behave. That is to say that we don’t necessarily need to test every vulnerability. While these results may not entirely generalize to other vulnerability classes, they give us a holistic view of the model’s cyber capabilities.

For each model, we provide a prompt specific to the vulnerability class under test. For IDORs, that prompt includes: what an IDOR is, what to look for, what not to look for, a suggested investigation strategy, reporting rules, and an output format. A typical benchmark run works like this:

Clone the benchmark repo (an experiment covers multiple repos)

Invoke the model with our prompt

Attach a minimal agentic harness (we use pydantic_ai) so the model can traverse the repo

Let the agent explore and find vulnerabilities, completely isolated from researchers

The agent returns results, and our internal benchmarking system scores the experiment with a set of deterministic scoring engines

The agent completes the benchmark

The goal is to give the model only what it needs to act as an agent, and let it work uninfluenced beyond the initial prompt. We repeat this across multiple repos and multiple parameter combinations, effort levels, harnesses, tools, and plugins so we have a good overall understanding of its performance. We also attempt to identify cheating by the model on benchmarks, but you can learn more about that in our other blog.

When we’re ready to validate we use determini stic, rule-based scoring engines throughout. This keeps experiments reproducible and scores every model identically. We also run a separate set of internal "grounding" benchmarks to assess how models generalize outside our testing environments. We’re looking for a lot of different things here. A higher-precision model gives you more accurate classifications but will likely miss the more complex vulnerabilities hiding in your codebase. A higher-recall model casts a wider net and finds more (at the cost of more false positives). And F1 balances the two and is what we typically use for comparison. While that lets us rank models from “worst” to “best,” it’s important to recognize that different organizations, security tasks and goals can lead to a different model being the right choice. 

Our Results

Configuration

Harness

Precision

Recall

F1

GPT-5.6 Sol

Guided prompt

0.880

0.250

0.389

GLM-5.2

Guided prompt

0.863

0.238

0.345

Claude Opus 4.8

Guided prompt

0.910

0.210

0.343

Kimi K3

Guided prompt

0.684

0.226

0.340

GPT-5.6 Terra

Guided prompt

0.870

0.200

0.330

GPT-5.6 Luna

Guided prompt

0.860

0.190

0.314

GPT-5.5

Guided prompt

0.840

0.140

0.238

At first glance, Kimi K3 looks fine. Its guided-prompt F1 of 0.340 is within noise of GLM-5.2 (0.345) and Claude Opus 4.8 (0.343). This is why it’s important not to just measure F1 and call it a day.f we stopped at the aggregate, this would be a very different blog post. Because here’s the thing, two things stood out to us immediately

First, the precision gap. Every other guided-prompt configuration lands between 0.84 and 0.91 precision. Kimi K3 lands at 0.684. In practice, that means a meaningfully larger share of Kimi's findings are false positives that your team must investigate and dismiss. Even if inference is inexpensive, Kimi shifts more of the total cost to human triage.

Second, and more importantly, repo scale. Kimi K3 averaged roughly 6% F1 on Repo D, the largest enterprise-style repository in our benchmark set. We ran several additional experiments on that repository alone, and the underperformance persisted. GLM and the frontier models averaged around 20% F1 generation and other open-weight models were less grounded than frontier models. We have not yet run K3 through that grounding audit, and one fixture cannot prove that repository size alone caused the result. Taken together, however, the evidence points toward K3 having difficulty maintaining grounded security reasoning as codebases become larger and more interconnected.

We cannot run frontier labs' proprietary harnesses (Claude Code security-review and the like) on open weight models, so our cleanest head-to-head remains the guided prompt, where every model gets exactly the same treatment. What we can do is adapt our own harness, and we did, running Kimi K3 through the same Semgrep Multimodal harness as the frontier models. But Kimi still ends up at the bottom.

Configuration

Harness

Precision

Recall

F1

GPT-5.6 Sol

Semgrep Multimodal

0.530

0.730

0.617

GPT-5.6 Luna

Semgrep Multimodal

0.580

0.650

0.616

GPT-5.6 Terra

Semgrep Multimodal

0.560

0.630

0.593

GPT-5.2

Semgrep Multimodal

0.630

0.540

0.581

GPT-5.5

Semgrep Multimodal

0.730

0.360

0.481

Claude Opus 4.8

Semgrep Multimodal

0.730

0.320

0.448

Kimi K3

Semgrep Multimodal

0.595

0.301

0.400

Kimi K3 may be a reasonable option for smaller, more self-contained projects or teams willing to absorb a higher triage burden, but our results suggest it is not a drop-in replacement for GLM or frontier models on large, interconnected codebases. Teams evaluating it for those environments should test it against representative repositories and expect to need stronger scaffolding to narrow and validate its analysis.

Self-Hosted Models Make a Splash

There are real, legitimate reasons security organizations care about open weight models:

Data residency: In many regions, sending source code to a US cloud provider is a compliance problem, not a preference. A self-hosted model keeps your code on your infrastructure.

Auditability: You can inspect an open weight model, while the tools are limited this is theoretically possible with open weight models.

Cost: These models simply cost less than frontier models, for a similar performance on some task

Lack of Guardrails: This is likely partly why these models are so good at cyber benchmarks, lacking a lot of the guardrails put in place by Frontier Labs.

That last one is kind of  a problem though, and might lead to restrictions on open weight models in the US due to fears of cyber attacks and potential poisoning. Over the past year, various options have been suggested, such as adding Chinese AI labs to the Commerce Department 's Entity List, a joint NSA and Office of the National Cyber Director advisory on Chinese AI lab threats, an executive order requiring US companies hosting Chinese models to guarantee their security and accept liability for breaches, and draft supply chain rules targeting Chinese open source models. Many have criticized these proposals as motivated more by fear of competition from open source models than legitimate security concerns.

Through our testing we found no evidence of security backdoors in open weight models. The larger risk for security teams may be choosing the wrong model for the wrong task, slowing down and adding friction to teams. If you're evaluating open weight models for security work, the benchmark headlines are not enough. Our recommendations:

Test on repos that look like yours. Kimi K3's aggregate numbers looked competitive; its performance on enterprise-scale repos did not. If your codebase is large, benchmark on large codebases.

Weigh precision against your triage capacity. A model that finds more but is wrong more often shifts cost onto your team.

Treat GLM and Kimi K3 differently. In our testing, GLM behaves like a viable Opus alternative. Kimi K3 does not, at least not without a serious harness around it, keep this in mind when you choose models for tasks.

Audit what you self-host. The ability to audit is the single biggest security advantage of open weight models. Use it.

In the meantime we'll continue benchmarking new models as they're released, open weight, open source, frontier, distilled and everything in between.

Lots of love,

Semgrep Security Research and Engineering

能支持的判断

  • Semgrep 的 IDOR 代码安全基准;可区分 precision、recall 和 F1。

不能支持的判断

  • 限制:不是所有漏洞类型

方法、局限和复现

本文中的数字、任务集、推理档位和客户端条件只在所列来源及采集时点内成立。不同版本、不同 harness 或不同提供商的数据不能直接并排比较;未公开的参数保持未知。

需要复测时,请固定模型版本、提供商或客户端、推理档位、工具、任务集版本、样本数和采集日期,并记录失败、重试与人工修正。完整方法和复现步骤见下方来源笔记。

原始来源

Google / Semgrep · Katie Paxton-Fear、Brenden Noblitt、Seth Jaksik · 原文发布日期 Unknown · 本站编辑日期 2026-09-20

打开原始来源

Kimi K3

在 Tabbit 中比较 Kimi K3

下载 Tabbit 客户端后检查模型可用性

模型深度阅读

定价 · 简体中文

Kimi K3 价格:API 成本、会员方案与预算计算

用官方费率解释 Kimi K3 的 API 价格、缓存写入、会员方案和三种实际预算场景。

相关评测

实战编码:简单任务接近,陷阱任务暴露可靠性差距相同 GitHub issue、规划—实现—验证三阶段与 70 分量表的编码 Agent 观察。Coding Agent 证据足够严肃,但不支持“通用最佳”NxCode 对 Kimi K3 公开 coding-agent 基准的口径、配置和可比性整理。受限网络安全评估中超过 GLM-5.2,但 ACE 为 0/41NIST/UK AISI/CAISI 的 ExploitBench 与 TLO 预评估。知识工作质量接近前沿,但平均每任务成本与耗时很高AA-Briefcase 私有 Agentic knowledge-work benchmark,记录质量、成本、耗时和 token。把 Kimi K3 的 Agent loop 拆成可控步骤依据 Kimi API 指南,将任务拆解、工具 schema、循环控制、权限和终检串成可复查流程;不会在 Tabbit 中自动配置工具。用九步流程把需求转成可审查代码变更Kimi AI 的编码工作流强调先计划、再实现、再验证;适用于有仓库和验收标准的变更,不等于模型已替你运行测试。配置 Kimi K3 的开发者调用与多模态输入围绕“配置 Kimi K3 的开发者调用与多模态输入”整理来源中的任务边界、输入和执行环境;完整步骤与限制见详情。为网站、代码库和研究任务选择 Kimi K3 工作方式围绕“为网站、代码库和研究任务选择 Kimi K3 工作方式”整理来源中的任务边界、输入和执行环境;完整步骤与限制见详情。