Kimi K3 is Moonshot AI's flagship long-context model for coding, research, knowledge work, visual reasoning, and agentic workflows. It is positioned as an open 3T-class frontier model: 2.8 trillion total parameters, native vision capability, and a 1,048,576-token context window.
This guide combines Moonshot's official Kimi K3 technical blog with Kie.ai's pricing analysis to explain what Kimi K3 is, how it works, where it is available, what it costs, and when it is worth using.
What Is Kimi K3?
Kimi K3 is Moonshot AI's most capable Kimi model at launch. In Moonshot's own description, it is built for "frontier intelligence" across long-horizon coding, knowledge work, and reasoning.
The most important characteristics are:
| Area | Kimi K3 detail |
|---|---|
| Scale | 2.8T total parameters |
| Architecture | MoE with Kimi Delta Attention and Attention Residuals |
| Context | 1,048,576 tokens |
| Modality | Native vision capabilities |
| Reasoning mode | Max thinking effort by default at launch |
| Availability | Kimi.com, Kimi Work, Kimi Code, and Kimi API |
| API model ID | kimi-k3 |
Moonshot says Kimi K3 still trails the strongest proprietary models in overall user experience, but reports frontier-level performance across its internal evaluation suite and strong results against other tested models.
Why Kimi K3 Matters
Kimi K3 matters because it pushes three ideas at once:
- Open-scale model size: Moonshot calls Kimi K3 the first open model to reach the 2.8T / 3T-class range.
- Practical long context: the 1M-token context window is not just for document stuffing; it targets coding agents, research agents, and long-running stateful workflows.
- Agentic product integration: Kimi K3 is not only an API model. It ships into Kimi.com, Kimi Work, Kimi Code, and enterprise workflows.
For users, this means Kimi K3 is less like a simple chat model and more like an agent model designed to keep working across code, documents, tools, visuals, and long chains of actions.
Architecture: KDA, Attention Residuals, and Sparse MoE
Kimi K3 is built on several architectural updates that Moonshot highlights as central to its scaling efficiency.
Kimi Delta Attention
Kimi Delta Attention, or KDA, is the attention mechanism behind Kimi K3's long-context design. Moonshot frames it as part of the model's foundation for handling information flow across very long sequences.
For users, the practical implication is that Kimi K3 is designed to keep more prior context available during long coding, research, and knowledge-work sessions.
Attention Residuals
Attention Residuals, or AttnRes, are described as a way to improve how representations move across model depth. In practice, this matters for deep reasoning and long-horizon tasks where the model must retrieve and reuse earlier internal state rather than treating each generation step as a shallow continuation.
Stable LatentMoE
Kimi K3 uses a Mixture-of-Experts design at very large scale. Moonshot says Kimi K3 effectively activates 16 of 896 experts with Stable LatentMoE.
The important point is that Kimi K3 does not activate the whole 2.8T model for every token. Like other sparse MoE systems, it routes tokens through selected expert paths, which helps make a very large model more practical to train and serve.
Scaling Efficiency
Moonshot reports that the combination of architectural changes, training recipes, and data recipes gives Kimi K3 about a 2.5x improvement in overall scaling efficiency compared with Kimi K2.
That does not mean every task is 2.5x better. It means Moonshot believes the model converts additional compute into capability more effectively than its predecessor.
Native Vision and Multimodal Work
Kimi K3 is not just a text model with image input bolted on. Moonshot describes it as having native vision capabilities, and the official examples emphasize work that blends code, visuals, screenshots, and iteration.
This matters for tasks such as:
- Frontend implementation from screenshots
- Game development with visual feedback
- CAD-style spatial reasoning
- Chart and dashboard generation
- Presentation and infographic design
- Visual QA during code generation
For a Chatmax AI workflow, that makes Kimi K3 relevant to both chat and creation flows: a user can reason over documents, inspect a visual, improve a UI, then continue into implementation.
Coding Capabilities
Kimi K3's strongest public positioning is long-horizon coding. Moonshot says it can sustain long engineering sessions, navigate large repositories, and orchestrate terminal tools with limited human oversight.
The official blog highlights several coding-oriented examples:
- Kernel optimization across GPU workloads
- Building a compact Triton-like compiler called MiniTriton
- Creating playable browser-based 3D experiences
- Using visual feedback loops for frontend and game development
- Designing and verifying a small chip in an autonomous run
- Reproducing scientific computational pipelines from literature
The point is not that every user should ask Kimi K3 to build a compiler or chip. The point is that Kimi K3 is designed for tasks where the model must plan, execute, inspect outputs, recover from errors, and continue for many steps.
Kimi K3 for Research and Knowledge Work
Kimi K3 is also positioned for end-to-end knowledge work. The official examples include consulting-style research, scientific analysis, interactive reports, dashboards, and visual presentations.
Moonshot's Kimi Work examples show Kimi K3 producing outputs such as:
- Interactive industry research websites
- Consulting reports with charts and timelines
- Scientific event analysis with concurrent subagents
- Editable heatmaps and annual-report-style visuals
- Persistent widgets and dashboards organized around a project
This is where the 1M-token window becomes useful. A knowledge-work agent often has to ingest many PDFs, reports, notes, charts, web pages, prior messages, and project constraints. Kimi K3 is built for that kind of dense workspace context.
Video Editing and Motion Design
One of the more interesting parts of the official Kimi K3 blog is its emphasis on motion design and video editing. Moonshot says Kimi K3 can understand text, images, and video in the same multimodal architecture.
The examples include creating a motion-graphics explainer of its own architecture and editing a teaser video from dozens of source clips. For normal users, the relevant takeaway is that Kimi K3 can reason about visual sequence, timing, clip selection, transitions, and revision loops.
This does not replace a full video production workflow, but it does make Kimi K3 relevant for planning, rough cuts, storyboard logic, and creative direction.
Availability: Where You Can Use Kimi K3
Moonshot lists several official ways to access Kimi K3:
| Access path | Best for |
|---|---|
| Kimi.com | General chat and testing |
| Kimi mobile app | Casual use and mobile access |
| Kimi Work | Knowledge work, documents, dashboards, research |
| Kimi Code | Terminal and IDE-oriented coding workflows |
| Kimi API | Developer integration and production systems |
| Kimi Enterprise | Organization-level privacy and member management |
For builders, the Kimi API is the important route. For evaluation, Kimi.com and Kimi Code are useful places to test behavior before building a production integration.
Quick Summary
As reported by Kie.ai's Kimi K3 pricing guide, the published Kimi K3 API rate card is:
| Item | Price |
|---|---|
| Input tokens, cache miss | $3.00 per 1M tokens |
| Input tokens, cache hit | $0.30 per 1M tokens |
| Output tokens | $15.00 per 1M tokens |
| Context window | 1,048,576 tokens |
| Model ID | kimi-k3 |
| Default max output | 131,072 tokens |
The important part is not only the $3 / $15 headline. The cache-hit input rate is 10x cheaper than a cold input request, so repeated agent sessions, long system prompts, and stable document prefixes can materially reduce cost.
Moonshot's official blog also states that the Kimi API achieves a cache hit rate above 90% in coding workloads, powered by its Mooncake disaggregated inference architecture. That is a major reason Kimi K3 can be priced competitively despite its model scale and long context.
What Kimi K3 Costs in Practice
Per-token prices are hard to reason about until they are converted into real workloads. These example totals use the published $3/M input, $0.30/M cached input, and $15/M output rates.
| Scenario | Input | Output | Cache | Estimated cost |
|---|---|---|---|---|
| Short chat turn | 2,000 | 500 | Miss | $0.014 |
| Short chat turn with warm cache | 2,000 | 500 | Hit | $0.008 |
| Long document Q&A | 200,000 | 2,000 | Miss | $0.63 |
| Long document Q&A with cached prefix | 200,000 | 2,000 | Hit | $0.09 |
| Full 1M-context extraction | 1,048,576 | 5,000 | Miss | ~$3.22 |
| Coding agent with heavy output | 20,000 | 40,000 | Hit | ~$0.61 |
The main takeaway: Kimi K3 becomes much more attractive when the same large context can be reused. If every request is cold and output-heavy, the output side quickly dominates the bill.
Pricing Formula
For API budgeting, use this simple formula:
total cost =
uncached_input_tokens / 1,000,000 * $3.00
+ cached_input_tokens / 1,000,000 * $0.30
+ output_tokens / 1,000,000 * $15.00This formula is more useful than looking at headline prices because Kimi K3 workloads can vary dramatically. A 200K-token document task with a cached prefix can be cheap; a short prompt that generates a very long answer can become output-heavy.
Output Tokens Are the Hidden Cost
Kimi K3's output price is five times the cold input price and fifty times the cached input price. That means long reasoning traces, verbose answers, tool logs, code diffs, and repeated retries can dominate the cost.
For production use, consider:
- Setting a realistic
max_completion_tokens - Asking for concise intermediate summaries
- Saving full traces only when needed
- Using structured outputs for extraction tasks
- Splitting exploratory reasoning from final user-facing answers
- Measuring cost per finished task, not only cost per request
Why Cache Hits Matter
Kimi K3's cache-hit rate is the biggest pricing lever. A stable prompt prefix, long repository context, repeated documentation block, or persistent agent instruction can move input tokens from $3.00/M to $0.30/M.
That matters for common AI workflows:
- Coding agents that keep the same repository map in context
- Research assistants that repeatedly query the same source pack
- Customer-support agents with a stable policy and knowledge prefix
- Long-document Q&A where only the question changes
- Multi-step planning sessions where the same instructions remain fixed
The practical rule is: avoid changing the prefix unnecessarily. If the app constantly rewrites system instructions or swaps model settings mid-session, cache efficiency can drop.
How to Design Prompts for Better Cache Efficiency
If you are building an app around Kimi K3, design prompts so the stable portion comes first and changes as little as possible.
A good request layout usually looks like this:
1. Stable system instructions
2. Stable tool policy or agent rules
3. Stable project documentation or repository map
4. Stable source documents
5. Changing user request
6. Short task-specific constraintsAvoid constantly inserting timestamps, random IDs, reordered documents, or user-specific metadata into the stable prefix. Small changes near the front of a request can reduce cache reuse.
1M Context Does Not Mean Every Call Should Use 1M Tokens
Kimi K3's 1,048,576-token window is useful for large codebases, legal packs, research corpora, and long multi-agent state. But a huge context window is not a free pass to send everything every time.
Use the full window when the task genuinely needs it:
- Cross-file refactors across many source files
- Long legal, finance, or policy review
- Multi-paper research synthesis
- Agent plans with extensive prior state
- Large debugging sessions with logs, code, and previous attempts
For simpler prompts, smaller context is cheaper, faster, and easier to control.
When the 1M Context Window Helps Most
Kimi K3's long context is most valuable when the answer depends on relationships across many pieces of information.
Good examples:
- Understanding a monorepo before making a cross-cutting change
- Comparing multiple policy documents for contradictions
- Reviewing financial reports across quarters
- Reading a large litigation or due-diligence pack
- Keeping a long coding-agent session alive
- Synthesizing many scientific papers
- Producing dashboards or slides from a large evidence pack
Weak examples:
- Simple rewriting
- Short Q&A
- Basic translation
- Small snippets of code
- Single-image captioning
- Quick social copy
For small tasks, Kimi K3 may still work well, but the long-context advantage is not doing much.
Kimi K3 API vs Consumer Access
Kie.ai reports that Kimi K3 is available through Moonshot's own API and consumer Kimi surfaces, with chat access available in the Kimi app and web product under normal usage limits.
For builders, the API is the relevant path because it gives you predictable integration and usage-based billing. For casual testing, the consumer app is the cheapest way to explore the model before committing to a production workflow.
Kimi Code and Agent Workflows
Kimi Code is one of the clearest product surfaces for Kimi K3. The official blog says users can run Kimi Code in a terminal and select Kimi K3 with the /model command.
This matters because a coding agent is one of the best environments for Kimi K3's strengths:
- Large context from repository files
- Repeated instructions and stable project state
- Terminal tool execution
- Long debugging loops
- Visual inspection for frontend work
- Code generation plus validation
The same properties make it risky if the agent is unconstrained. Kimi K3 can be proactive, so production agent setups should define explicit boundaries around file changes, terminal commands, deployment, and external calls.
Subscription and Top-Up Notes
The referenced pricing guide also notes Chinese-market prepaid credit plans and launch top-up bonuses. These are useful for understanding early access economics, but they should not be treated as permanent pricing.
Before budgeting a production deployment, verify:
- Current Moonshot API rates
- Whether launch bonuses are still active
- Whether Kimi Code plans expose the context size you need
- Whether K3 Swarm Max has separate billing
- Whether web search, vision input, or batch processing has separate charges
Deployment and Self-Hosting Reality
Kimi K3 is open in intent, and Moonshot says full model weights will be released by July 27, 2026. But self-hosting a 2.8T-parameter MoE model is not a practical cost-saving path for most teams.
Moonshot's own infrastructure notes point to serious serving requirements:
- Quantization-aware training from the SFT stage
- MXFP4 weights with MXFP8 activations
- Large expert-parallel training and serving concerns
- vLLM-related work for KDA prefill cache
- Recommended deployment on supernode configurations with 64 or more accelerators
For most product teams, the official API is the realistic path. Open weights are valuable for research, infrastructure partners, and frontier model transparency, but they do not automatically make deployment cheap.
How Kimi K3 Compares to Other Models
Kimi K3 sits in a different pricing band from older Kimi K2 models. It is more expensive than prior Kimi generations, but it also targets larger context, deeper reasoning, and agent workflows.
Compared with premium closed models, Kimi K3's headline input and output rates can look competitive. But the real comparison depends on:
- Average output length
- Reasoning verbosity
- Cache-hit ratio
- Tool-use cost
- Failure and retry rate
- Whether the task benefits from 1M context
A cheaper per-token model is not automatically cheaper per finished task if it produces longer reasoning traces or needs more retries.
Kimi K3 vs Kimi K2
Kimi K2 was already known for large-scale open-model ambition, but Kimi K3 raises the ceiling in several ways:
| Dimension | Kimi K2 | Kimi K3 |
|---|---|---|
| Positioning | Earlier Kimi flagship family | New most capable Kimi model |
| Scale | Smaller than K3 | 2.8T total parameters |
| Context | Shorter than K3 | 1,048,576 tokens |
| Architecture | Earlier generation | KDA, AttnRes, Stable LatentMoE |
| Focus | Chat, coding, reasoning | Long-horizon coding, knowledge work, vision, agents |
The upgrade is not just parameter count. Kimi K3 is a broader agentic model with stronger emphasis on long sessions, multimodal reasoning, and tool-oriented execution.
Kimi K3 vs GPT-5.6 Sol and Claude Fable 5
Moonshot explicitly says Kimi K3 still trails the most powerful proprietary models in overall performance and user experience, including GPT-5.6 Sol and Claude Fable 5. That caveat matters.
Kimi K3's strongest argument is not "best at everything." Its strongest argument is:
- Very large context
- Open-model positioning
- Competitive frontier performance
- Native multimodal capability
- Strong coding and knowledge-work orientation
- API pricing with aggressive cache economics
If you want the most polished proprietary assistant experience, GPT-5.6 Sol or Claude Fable 5 may still be the safer default. If you want an open frontier-scale model with 1M context and strong agent capabilities, Kimi K3 deserves evaluation.
Best Use Cases for Kimi K3
Kimi K3 makes the most sense when the value of long context outweighs the cost:
- Large codebase analysis
- Multi-document research
- Long policy or legal review
- Agentic coding loops
- Repeated workflows with stable prompt prefixes
- Context-heavy assistants where cache hits are likely
It is less compelling for very short, one-off tasks where a smaller or faster model can produce the answer at lower latency and cost.
Practical Prompting Tips
Use Kimi K3 like a long-horizon agent, not like a short autocomplete model.
Good prompt patterns:
- Give it the goal, constraints, and stopping condition.
- Tell it what files, sources, or documents matter most.
- Ask it to summarize its plan before making large changes.
- Use explicit boundaries for risky actions.
- Ask for compact final answers when cost matters.
- Reuse stable context across turns instead of rebuilding prompts.
For coding:
- Include repository structure before individual files.
- Put coding standards and test commands in the stable prefix.
- Ask for small verified patches before broad rewrites.
- Keep the same session alive when cache and context matter.
For research:
- Group sources by reliability and topic.
- Ask for citations or source mapping.
- Separate extraction, synthesis, and recommendation steps.
- Use tables when comparing many documents.
For dashboards and presentations:
- Describe the target audience and decision the artifact must support.
- Provide data definitions before asking for visuals.
- Ask for a first outline before requesting polished output.
Open Questions
Some details are still worth monitoring:
- Whether K3 Swarm Max uses the same pricing model
- Whether batch discounts become available for K3
- How vision inputs are billed
- Whether web-search tool pricing changes
- Whether open weights materially reduce real deployment cost
- How long promotional top-up rates remain available
These details can change the economics for production systems.
Known Limitations
Moonshot lists several limitations that are important for production users.
Thinking-history sensitivity
Kimi K3 was trained with preserved thinking-history behavior. Moonshot warns that quality can become unstable if an agent harness does not pass historical thinking content correctly or if a session switches to Kimi K3 in the middle.
The practical advice is simple: start Kimi K3 sessions cleanly and use compatible harnesses such as Kimi Code when possible.
Excessive proactiveness
Kimi K3 is trained for long-horizon, difficult tasks. That can make it overly proactive when requirements are ambiguous. It may make decisions on the user's behalf unless the system prompt clearly defines boundaries.
For business software, this means you should explicitly constrain:
- What files the agent can edit
- Whether it can run terminal commands
- Whether it can call external tools
- Whether it can deploy, publish, or delete
- When it must ask for human confirmation
Proprietary-model user experience gap
Moonshot also notes that Kimi K3 remains behind the strongest proprietary systems in overall user experience. This is a useful reminder: model selection should be based on actual workflow tests, not only benchmark headlines.
FAQ
How much does Kimi K3 cost?
The published API pricing referenced by Kie.ai is $3.00 per 1M input tokens on a cache miss, $0.30 per 1M input tokens on a cache hit, and $15.00 per 1M output tokens.
Is Kimi K3 free?
Kimi K3 can be tested through consumer Kimi surfaces under normal usage limits, but production API use is pay-as-you-go.
What is the cheapest way to use Kimi K3?
For API workloads, the cheapest pattern is to keep stable prompt prefixes and maximize cache hits. For simple testing, the consumer Kimi app is the lowest-friction entry point.
How much does a full 1M-token Kimi K3 request cost?
A cold 1,048,576-token input costs about $3.15 before output. With a cache hit, the same input is about $0.31. Output is billed separately at $15 per 1M tokens.
Is Kimi K3 good for coding agents?
Yes, especially when the agent repeatedly uses the same repository context or system instructions. The large context window and cache-hit pricing can fit long coding loops, but output-heavy reasoning still needs budget control.
Should every app switch to Kimi K3?
No. Kimi K3 is best for long-context and agentic workloads. Short support replies, simple rewriting, or low-latency tasks may be better served by smaller models.
What makes Kimi K3 different from Kimi K2?
Kimi K3 is larger, uses a new architecture, supports a 1M-token context window, includes native vision capability, and is positioned more strongly around long-horizon coding, knowledge work, and agent workflows.
Is Kimi K3 open source?
Moonshot describes Kimi K3 as an open 3T-class model and says full model weights will be released by July 27, 2026. For most users, API access will still be the practical way to use it because the infrastructure requirements are substantial.
Does Kimi K3 support vision?
Yes. Moonshot describes Kimi K3 as having native vision capabilities and shows examples involving screenshots, visual reasoning, game development, dashboards, presentations, and video editing.
Is Kimi K3 good for long documents?
Yes. The 1,048,576-token context window makes it suitable for large document packs, research corpora, codebases, and long agent sessions. The best results still require careful source organization and output constraints.
What is the main risk when using Kimi K3?
The main practical risks are output cost, over-proactive agent behavior, and instability if the harness does not preserve context in the way Kimi K3 expects.
Bottom Line
Kimi K3 is one of the most ambitious open frontier-model launches to date: 2.8T parameters, native vision, 1M context, sparse MoE architecture, and strong focus on coding agents and knowledge work.
Its pricing is attractive when you use the model as intended: long context, stable prefixes, and repeated workflows that benefit from cache hits. The risk is output-heavy reasoning. If responses become very long, the $15/M output rate can dominate total cost.
For teams building model directories, AI chat products, coding agents, or document-heavy assistants, Kimi K3 is worth testing. Just measure cost per completed task, not only cost per token.

