01 / MODEL
Thinking is always on
Moonshot documents K3 as an always-thinking model. The API accepts reasoning_effort low, high, or max. A long reply is not automatically a broken preset.
SETUP BEFORE PRESETS
A working connection starts with four details: the live provider, the model alias, the Chat Completions path, and how thinking history is handled. Check those before tuning a character preset.
Official alias: kimi-k3 · Example base URL: https://api.moonshot.ai/v1 · K3 always has thinking enabled.

WHAT THE SIGNALS MEAN
Kimi K3 can be a good writing model and still expose a bad endpoint or an incomplete reasoning message. These three checks keep the diagnosis honest.
01 / MODEL
Moonshot documents K3 as an always-thinking model. The API accepts reasoning_effort low, high, or max. A long reply is not automatically a broken preset.
02 / PROVIDER
The official example uses https://api.moonshot.ai/v1 and model kimi-k3. Proxies can expose different aliases, quotas, prices, or parameter support.
03 / REPORTS
Public X and Reddit discussions mention slow responses, 429s, overthinking, and reasoning blocks. They describe experiences, not a service-level promise.
THE FIVE-CHECK SETUP
Use a small test chat first. Change one setting, send one turn, and keep a note of what changed.
K3 API facts come from the official Kimi guide. Provider dashboards and SillyTavern extension behavior can change independently.
Choose Moonshot or a provider that currently lists Kimi K3. Confirm billing, rate limits, the live model slug, and whether the route is OpenAI compatible.
✓ You can point to the provider entry that says Kimi K3 or kimi-k3.
The official Quickstart sends requests to /chat/completions. In SillyTavern, use the Chat Completions path rather than a legacy text-completion template.
✓ The connection test reaches /v1/chat/completions without a template error.
Use one character, one scene, and a small context window. State voice and boundaries once. Contradictory patches make it harder to tell whether the API works.
✓ The first reply has the expected role and format before any preset is added.
K3 streaming can return reasoning_content separately from content. Make sure the frontend handles reasoning blocks and sends the complete assistant message on the next turn.
✓ The next request does not silently drop the assistant reasoning or final message.
Prefill and preserved-thinking extensions can help with a specific RP behavior, but they are community tooling, not a Kimi requirement. Change one layer and keep a rollback copy.
✓ You can name the one change that improved or worsened the reply.
SYMPTOM FIRST
Use the narrowest check that matches the symptom. Do not rotate through random presets while an auth or endpoint error remains.
| What you see | Likely layer | Next check |
|---|---|---|
| 401 or "invalid key" | Key or account | Create a fresh key in the provider console, confirm the account is unlocked, and remove stray spaces. |
| 404 or model not found | Base URL or alias | Compare the live provider model list with kimi-k3. Keep /v1 in the base URL and let the connector add /chat/completions. |
| 429, timeout, or “too busy” | Quota or rate limit | Check concurrency, RPM/TPM, account tier, and provider status. Retry only after confirming the request is not looping. |
| Empty reply or broken reasoning block | Response mapping | Inspect reasoning_content versus content, enable the connector reasoning option, and preserve the complete assistant message. |
| Long thoughts, drift, or repetition | Prompt and history | Lower reasoning_effort when the provider supports it, shorten the instruction stack, and test a clean chat before changing prefill. |
A SHORTER PATH
If the immediate job is reading a character sheet, wiki, or writing brief while chatting, Tabbit removes the first endpoint setup loop. Check the live picker for access and availability.
Open the character sheet, lore page, or reference document in a tab. Tabbit can use visible pages, screenshots, and files as context.
✓ The source stays visible while you ask the question.

Select Kimi-K3 from the current Chat or new-tab model picker. The model roster can change, so treat the live UI as the source of truth.
✓ The picker shows the model you intend to use.

Use the sidebar for a focused question or compare model replies in one view. Tabbit is a browser chat route, not a replacement for ST cards, lorebooks, or group chat.
✓ Your prompt keeps the page or file context in the same workflow.

SETUP FAQ
The official Quickstart uses model kimi-k3 and the example base URL https://api.moonshot.ai/v1. Verify the live provider list because aliases and availability can change.
Use an OpenAI-compatible Chat Completions connection for the official Moonshot route. A text-completion template is a different request shape.
K3 always has thinking enabled in the Moonshot API. Where supported, try a lower reasoning_effort, keep instructions short, and confirm the full assistant message is preserved.
No. Start with the clean connection first. Add community preset or prefill guidance only when you can describe the behavior you want to change.
Treat 429 as a quota or rate-limit signal. Check account tier, concurrency, RPM/TPM, provider status, and whether SillyTavern is retrying repeatedly.
No. Tabbit is useful for using Kimi K3 beside a live page or file. SillyTavern remains the tool for its character cards, lorebooks, extensions, and group chat.
Keep SillyTavern for its character workflow. For a lower-friction first chat, open Tabbit, choose Kimi-K3, and bring the page context with you.
Available for macOS and Windows. Model access and quotas may vary by edition and plan.