Tabbit
活动资源博客模型
Tabbit LogoTabbit

Tabbit — 为你工作的 AI 浏览器

主题资源

  • AI Browser Resources
  • Agentic Browser Resources
  • Browser Downloads and Install Guides
  • Browser Comparisons
  • AI Browser Alternatives
  • Browser Productivity Resources

热门指南

  • AI Browser
  • Agentic Browser Download
  • Best AI Browser 2026: Top 9 Tested & Ranked
  • AI Browser Download
  • Free AI Browser
  • Best AI Browser 2026
  • AI Browser Comparison 2026
  • AI Browser for Windows
  • AI Browser for Mac
  • Chrome Alternative 2026

活动

  • 别装了,你在《牛来》里早有原型
  • Tabbit 妙招大赛
  • KPOP SBTI 饭圈人格测试
  • Tabbit 校园共创者计划
  • fifi 的论文文献妙招精选
  • 用户问卷

关于

  • Tabbit 博客
  • 媒体报道
测评
社区Kimi K3

Kimi K3 的实际使用感受与 benchmark 对照

原始来源

Reddit / r/LocalLLaMA

作者superSmitty9999 等 Reddit 用户

Tabbit 整理2026-08-20

查看原文

文章正文

跳到主要内容 How does Kimi k3 feel? Does it match up where it stands on benchmarks? : r/LocalLLaMA 在 Reddit 上投放广告 打开聊天 创建 创建帖子 打开收件箱 展开用户菜单 转载 转到“LocalLLaMA” r/LocalLLaMA • 1个月前 superSmitty9999 How does Kimi k3 feel? Does it match up where it stands on benchmarks? Kimi k3 的感受如何?它在基准测试中的表现是否与实际相符? Question | Help

I've seen the benchmarks, they supposedly right up there with Fable 5 and Sol 5.6. 我看过基准测试,据说它的表现与 Fable 5 和 Sol 5.6 不相上下。

Howerver I'm skeptical of benchmarks and the kimi k3 creators even mentioned the user experience isn't quite on par. 然而我对基准测试持怀疑态度,就连 Kimi K3 的开发者也提到,用户体验尚未达到理想水平。

How is K3 doing on coding? How is the personality? Does it's thinking and train of thought feel high quality or sort of insane? Common sense? K3 在编程方面表现如何?它的性格特征怎样?它的思考过程和思维脉络感觉质量高,还是有点疯狂?常识呢?

Trying to get a sense of the quality in real world use. 想了解在实际使用中其质量如何。

共享 kalshi_official • 已推广 2028 PRESIDENTIAL ELECTION ODDS: Check today's latest moves on Kalshi.com 了解更多信息 kalshi.com 排序方式: 评论区域 eli_pizza • 1个月前 前 1% 最受欢迎的评论者

Put a couple bucks in openrouter and you can try it. It’s good. 在 OpenRouter 里充几美元,你就可以试用了。很不错。

回复 共享 mo-powerbuilder • 1个月前

Is open router good? Open Router 好吗?

回复 共享 RecordingOk3922 • 1个月前

I personally like it 我个人很喜欢它

回复 共享 eli_pizza • 1个月前 前 1% 最受欢迎的评论者

I mean if you only want to use Kimi you can also just give them a few bucks and use it directly. But yes openrouter is good and lets you easily test pretty much any model through one endpoint. 我的意思是,如果你只想使用 Kimi,也可以直接给他们几美元,直接使用。但确实,OpenRouter 很好,让你可以通过一个端点轻松测试几乎任何模型。

回复 共享 SporksInjected • 1个月前 前 1% 最受欢迎的评论者

When the API is available 当 API 可用时

回复 共享 aboutthednm • 1个月前

Is it not live here? 它这里不是实时的吗?

https://openrouter.ai/moonshotai/kimi-k3

Do I misunderstand what you're asking? 我是不是误解了你的问题?

hidden2u • 1个月前 前 1% 最受欢迎的评论者 Regular-Anybody2645 • 1个月前

The whole world wants to use it and nobody else is hosting until open weights on the 27th at the least.

eli_pizza • 1个月前 前 1% 最受欢迎的评论者

It works though.

backyard_tractorbeam • 1个月前

K3 is available on opencode go now too. Or, it's at least in the price list, haven't tried it.

LMTLS5 • 1个月前

you can use it for free in their chat app too. it is indeed good

zeroccx • 27天前

Same. I didn’t really buy the hype until I used it myself. Honestly, if models like this keep coming out, the US should be more worried about staying competitive than trying to slow everyone else down by their nonsense restrictions

FBIFreezeNow • 1个月前

Feels like Fab Fable’s younger brother and Opus 4.8’s uncle and Sonnet’s mom. It certainly doesn’t feel like it’s related to gpt family though

jeekp • 1个月前

Thanksgiving is gonna be awkward this year

MeretrixDominum • 1个月前

I should put aside $1000 and have a group chat with them all at the same tims and see who I can seduce first

PhantomGaming27249 • 1个月前

Best ui model I have ever used and as good as 5.6 sol on the backend. Genuinely amazing model.

superSmitty9999 • 1个月前

Can you specify more about what you asked it and your use case? Thanks!

PhantomGaming27249 • 1个月前

GPU kernal optimization work and reworking a ui of a project I have been doing. Did excellent work on both tasks.

robkkni • 1个月前

Wow! That's some non-trivial work!

superSmitty9999 • 27天前

Awesome that they don't sandbag you on AI work lol

u/Kurrent-dot-io • 已推广 Capacitor is a free multiplayer shared coding session memory across Claude, Codex, Cursor, Pi, OpenCode and more. 查看更多内容 kurrent.io Admirable_Market2759 • 1个月前 前 1% 最受欢迎的评论者

It lives up to the hype

ShelZuuz • 1个月前

I gave it $100 today and told it to fix a Rust bug that I've just one-shot fixed with Fable on another machine in 15 minutes. Wanted to see what the equivalent cost would be for a direct comparison.

Then when I checked on it later it just ran itself out of credits and having gotten nowhere.

Exactly the same prompt as Fable. Basically described where an application output was drifting from the spec and told it to read the spec and look at the output then fix whatever is causing the application to not match the spec. It found the area in the spec, but then continued to read the spec (500 pages) for some unknown reason. And it made fixes which it was just guessing it without it being based on the spec.

I'm sure I can get it to write code if I hand it the correct C++ file and tell it what function to fix, like you would do with Haiku, but this model is supposed to compete with Fable head-on, not Haiku.

SentientPetriDish • 1个月前

same model that crashed the S&P 500, lmao

hellomistershifty • 1个月前

Yeah this was my experience with GLM-5.2. Just endless reading until the thinking hits the context limit, and repeat. Great for smaller stuff but hits a wall in a big codebase

ebolathrowawayy • 1个月前 martinerous • 1个月前 • 1个月前 编辑

Tried it on OpenRouter for a horror adventure roleplay. Surprising - it even speaks Latvian better than GLM (and Grok), but not as good as Gemini, GPT or Claude. Also, I liked how it builds the environment, feels less naive/cliche, when compared to Gemini, which I'm the most familiar with. Also, it followed my prompt instruction to make characters silent when they are alone. Gemini did not follow this, always wanting to "think aloud". Not good thing - coherence of characters actions somehow did not seem as good as Gemini Pro. Some stuff just did not make sense. It could be because of Latvian, and it might be better for English though.

And as it is smarter in general, it's also smarter with refusals - it picks up hints of things it doesn't like that potentially might lead to unethical behaviors. So, it will not play evil characters. I (naively) hope it's in the system prompt and not baked into the model.

And it's quite slow on OpenRouter. I imagine, the demand is huge now.

studansp • 1个月前

Just wanted to say sveiks 🇱🇻 - Latvian American and that's all I know.

martinerous • 1个月前

Thanks :) Yeah, we had quite many people leaving Latvia and finding their new home in the US during the tough war and USSR occupation times.

superSmitty9999 • 1个月前

If you're running it on api via OpenRouter, you control the system prompt right? So it's probably just censored.

martinerous • 1个月前

I think so, and that would be sad. But then, there might be hidden additional prompts or filters in-between the provider and OpenRouter. For example, Google have special content HarmCategory settings that are available in their direct API, and Gemini can become quite evil when all filters are turned down, but I suspect they are exposing Gemini on OpenRouter and elsewhere with more strict settings. Who knows, what Kimi providers are doing behind the scenes and if they have any filters...

u/RedditforBusiness • 已推广 Reach 443M+ high-intent audiences on Reddit. 注册 ads.reddit.com TimAndTimi • 1个月前

Mr. Dario responsed by making Fable 5 staying in Max amd Pro plans.

geteum • 1个月前

I gave a try today in a code optimization I have solved recently but sol failed. It failed as well hahaha but it was interesting what it suggested. For this problema I spent around 10 dollars with sol and 2 with Kimi.

The issue is a double loop with two operations ordered in way that lead to a huge memory consumption. Altering the order is enough to reduce the memory by a factor of ten. I dont know why but no LLM can figure this.

rpeck • 1个月前

Did you try hinting to them to think through the memory access pattern because you think that might be the problem? (aka is often the problem with inner loops like this)

is)

Temporary_Stick_6664 • 1个月前

you're too smart bro

real_serviceloom • 1个月前

So, I've been using it as my main model.. I do heavy backend work with Rust...

It definitely feels fable class. It is really good. It can also understand things very thoroughly, similar to feeble in that regard, where you can give it some sentence which makes sixty percent sense and can figure out what you mean exactly.

If I have to criticize it, maybe the only one is that it is really slow and it thinks quite a lot, but I think that's also because I haven't fully learned how to drive the model and that's something that comes with experience.

tomekza • 1个月前

"feeble" 😂

real_serviceloom • 1个月前

lmao just realized that typo. keeping it as i think its a good slip.

New_Alps_5655 • 1个月前

So far I'd say it's exactly where the benchmarks show. Only a hair behind Fable at a fraction of the price.

Cupakov • 1个月前

Haven’t tried it with coding tasks but with research it seems kind of hesitant to look stuff up. When reminded (or with a system prompt that emphasises search tools) it excels though, really impressed. I’m using it with GPT5.6 Sol as a fallback and it’s really hard for me to pick which produces better outputs.

ShamanJohnny • 1个月前

Every model has its strengths. The strengths of this model are the Design,UI/UX, and writing. For everything else i would just use Sol.

matrik • 1个月前 • 1个月前 编辑

I tried documenting a complex and badly written repo with it. It did an amazing job, far beyond opus, without finding any excuses or shortcuts. Downside is that it consumed the 5h budget in about 10 minutes, so I had to upgrade from moderato (cheapest) to allegretto (2nd cheapest). The whole job took ~30 minutes, and t/s was around 10-15. I believe they are experiencing overload due to the hype.

I'm very satisfied with the end result, but it was a bit more expensive than I expected.

Overall, it convinced me to test it as a daily driver. It would have been so much better if I can run it on my hardware.

Southern_Sun_2106 • 1个月前

If they don't screw around with model quality behind the API, like both OpenAI and Anthropic does, that in itself will be a reason to switch.

segmond • 1个月前 llama.cpp 前 1% 最受欢迎的评论者

send me some money and I'll try it and let you know.

superSmitty9999 • 1个月前 • 1个月前 编辑

What's your email and phone number I'll zelle it to you /s

NaiveDragonfruit • 1个月前

where are the model weights

Any-Conference1005 • 1个月前

27th

Ylsid • 1个月前

It's good but waffles too muc 前 • 1个月前 编辑

What's your email and phone number I'll zelle it to you /s

NaiveDragonfruit • 1个月前

where are the model weights

Any-Conference1005 • 1个月前

27th

Ylsid • 1个月前

It's good but waffles too much

jarec707 • 1个月前

Check out Ethan Mollick’s tweets: x or Bluesky. He provides a balanced view. TL;DR Kimi 3 is in a class below Fable and ChatGPT equivalent. Kimi made some significant errors in academic work. All in all a very good but not superb model.

PinEnvironmental6395 • 1个月前

Read it. He is coping extremely hard. 

jarec707 • 1个月前

I don't quite follow, say more? I've experienced him as a knowledgeable, experienced power user without an axe to grind.

DryWeb3875 • 1个月前

So is it Ethan Bollicks?

RadioactiveBread • 28天前

No idea what these people are on about when they say its "definitely Fable class". If you think it's Fable class you aren't doing tasks that require Fable.

DauntingPrawn • 27天前

Opus 4.8 without the kvetching.

ketosoy • 1个月前 前 1% 最受欢迎的评论者

In my minimal testing, it has lived up to the hype.

I tried it on an ascii art problem opus and gpt5.6 had both struggled with and it did great.

I also had it mock up a simple 3d education game “neon diner with tron vibes” and it nailed the vibe and the UI was legitimately good with 4 prompts.

NUMERIC__RIDDLE • 1个月前

I like it. Sometimes it says some out-of-pocket shit. Honestly, it finally feels like a model that can actually replace all of my Claude use for me. And that's on vibes.

创建于 2023年3月10日 公共 90万2.8万 用户标识 1llegi 社区书签 维基 Best LLMs Megathread  最佳 LLMs 综合帖 Best VLMs Megathread  最佳视觉语言模型合集帖 Best TTS/STT Models  最佳语音合成/语音识别模型 R/LOCALLLAMA 规则 1 Please search before asking   请先搜索再提问 2 Off-Topic Posts   非主题帖子 3 Low Effort Posts   低努力发帖 4 Limit Self-Promotion   限制自我推广 5 Follow Reddit's Content Policy   遵守 Reddit 的内容政策 SOCIALS  社交平台 AMA 版主 向版主发送消息 u/HOLUPREDICTIONS

Sorcerer Supreme   至尊法师 Yo  嘿 u/AskGrok Grok u/ArcaneThoughts u/Lissanro u/townofsalemfangay u/XMasterrrr

LocalLLaMA Home Server Final Boss 😎 @TheAhmadOsman u/rm-rf-rm

u/WithoutReason1729 u/No_Afternoon_4260

llama.cpp u/ttkciar

llama.cpp 查看所有版主 已安装的应用 Admin Tattler Bot Bouncer Reddit 规则 隐私政策 用户协议 你的隐私选择 辅助功能 Reddit, Inc. © 2026。保留所有权利。 折叠“导航” 创建社区 REDDIT 上的游戏 定制信息流 创建自定义信息流 最近访问 r/opencodeCLI r/chrome r/Notion r/todoist 社区 管理社区 资源 关于 Reddit 广告 开发者平台 Reddit Pro 测试版 帮助 博客 职业 新闻 Reddit 最佳 Reddit 规则 隐私政策 用户协议 你的隐私选择 辅助功能 Reddit, Inc. © 2026。保留所有权利。

Tabbit 小编提醒

本文是第三方资料导航。测试环境、模型版本和主观体验可能不同,请以原文为准。

Kimi K3

在 Tabbit 中使用并对比模型

Kimi K3

相关测评

媒体Google / Semgrep

Kimi K3 的代码安全测评:表面强劲,精度不足

媒体Google / MindStudio

Kimi K3 实战编码测评:真的像宣传的那么好吗?

媒体Google / Simon Willison

Kimi K3 与 pelican benchmark:我们还能学到什么

媒体Google / NxCode

Kimi K3 benchmark 详解:coding-agent 测评指南

Kimi K3

相关提示词

媒体Google / Business Compass LLC

Kimi K3 提示词工程指南

媒体Google / Together AI

Kimi K3:完整开发者指南

媒体Google / Kimi API Platform

Kimi 提示词最佳实践

媒体Google / Kimi API Platform

使用 Kimi K3 构建 Agent