Kimi K3 is Moonshot AI’s 2.8T-parameter flagship model for long-horizon agentic work, with native vision, always-on reasoning, and a 1M-token context window.
Kimi K3 is Moonshot AI's most capable flagship model to date, built for long-horizon agentic work in LM Studio Bionic. The 2.8-trillion-parameter model is the first open model to reach the 3T-parameter class.
low, high, and max effort levelsKimi K3 is designed for repository-scale coding, architecture work, and complex debugging. Its native vision support enables visual work such as frontend development: it can inspect screenshots, modify a codebase, and reason over the resulting interface.

Coding benchmark results as reported by Moonshot AI. Source.
For research and knowledge work, Kimi K3 can analyze and synthesize across large collections of files, extract findings, create reports and presentations, and sustain reasoning across long-running tasks.

General-agent and visual-agent benchmark results as reported by Moonshot AI. Source.
The following results are reproduced from Moonshot AI's Kimi K3: Open Frontier Intelligence. Scores, harnesses, comparison settings, and model names are reported by Moonshot AI; see the source article's footnotes for complete methodology and caveats.
| Category | Benchmark | Harness / note | Kimi K3 (max) | Claude Fable 5 (max, with fallback) | GPT 5.6 Sol (max) | Claude Opus 4.8 (max) | GPT 5.5 (xhigh) | GLM-5.2 (max) |
|---|---|---|---|---|---|---|---|---|
| Coding | DeepSWE | Kimi Code | 67.5 | 70.0 | 73.0 | 59.0 | 67.0 | 46.2 |
| Coding | Program Bench | Kimi Code | 77.8 | 76.8 | 77.6 | 71.9 | 70.8 | 63.7 |
| Coding | Terminal Bench 2.1 | Kimi Code | 88.3 | 84.6 | 88.8 | 84.6 | 83.4 | 82.7 |
| Coding | FrontierSWE | Dominance as of 26/7/16 · Kimi Code | 81.2 | 86.6 | 71.3 | 66.7 | 64.9 | 67.3 |
| Coding | SWE Marathon | Claude Code | 42.0 | 35.0 | 39.0 | 40.0 | 14.0 | 13.0 |
| Coding | PostTrain Bench | Claude Code | 36.6 | 41.4 | 34.6 | 34.1 | 28.4 | 34.3 |
| Coding | MLS Bench | Kimi Code | 48.3 | 49.9 | 46.2 | 42.8 | 35.5 | 40.4 |
| Coding | Kimi Code Bench 2.0 (Internal) | — | 72.9 | 76.9 | 64.8 | 71.7 | 69.0 | 64.2 |
| Agentic | GDPval-AA v2 (Elo-score) | — | 1668 | 1760 | 1748 | 1600 | 1494 | 1514 |
| Agentic | BrowseComp | — | 91.2 | 88.0 | 90.4 | 84.3 | 84.4 | — |
| Agentic | DeepSearchQA (f1-score) | — | 95.0 | 94.2 | — | 93.1 | — | — |
| Agentic | Toolathlon-Verified | — | 73.2 | 77.9 | 74.9 | 76.2 | 73.5 | 59.9 |
| Agentic | MCP Atlas | — | 84.2 | 84.7 | 83.6 | 83.6 | 82.8 | 82.6 |
| Agentic | Automation Bench | — | 30.8 | 29.1 | 29.7 | 27.2 | 22.7 | 12.9 |
| Agentic | Job Bench | — | 52.9 | 57.4 | 46.5 | 48.4 | 38.3 | 43.4 |
| Agentic | AA-Briefcase (Elo-score) | — | 1548 | 1583 | 1495 | 1354 | 1158 | 1260 |
| Agentic | APEX-Agents | — | 41.0 | 43.3 | 39.9 | 39.4 | 38.5 | 35.6 |
| Agentic | Office QA Pro | — | 63.3 | 69.9* | 63.2* | 63.9* | 60.9* | 41.4 |
| Agentic | SpreadsheetBench 2 | — | 34.8 | 34.7* | 32.4* | 31.55* | 29.05* | 28.12 |
| Agentic | DECK-Bench (Internal) | — | 73.5 | 73.0 | 74.7 | 66.9 | 68.2 | 68.6 |
| Reasoning & Knowledge | GPQA-Diamond | — | 93.5 | 92.6 | 94.1 | 91.0 | 93.5 | 91.2 |
| Reasoning & Knowledge | HLE-Full | — | 43.5 | 53.3 | 44.5 | 49.8* | 41.4* | — |
| Reasoning & Knowledge | HLE-Full w/ tools | — | 56.0 | 63.0 | 58.0 | 57.9* | 52.2* | — |
| Vision | MMMU-Pro | — | 81.6 | 81.2 | 83.0 | 78.9 | 81.2 | — |
| Vision | MMMU-Pro w/ python | — | 83.4 | 86.5 | 84.6 | 82.7 | 83.2 | — |
| Vision | CharXiv (RQ) | — | 84.8 | 88.9 | 84.6 | 80.5 | 84.1 | — |
| Vision | CharXiv (RQ) w/ python | — | 91.3 | 93.5 | 89.1 | 89.9 | 89.0 | — |
| Vision | MathVision | — | 94.3 | 94.8 | 95.8 | 86.7 | 92.2 | — |
| Vision | MathVision w/ python | — | 97.8 | 98.6 | 97.8 | 97.1 | 96.8 | — |
| Vision | BabyVision w/ python | — | 85.7 | 90.5 | 88.9 | 81.2 | 83.6 | — |
| Vision | ZeroBench_main (pass@5) | — | 23.0 | 23.0 | 17.0 | 17.0 | 22.0 | — |
| Vision | ZeroBench_main w/ python (pass@5) | — | 41.0 | 46.0 | 35.0 | 34.0 | 41.0 | — |
| Vision | WorldVQA ForceAnswer | — | 51.0 | 56.7 | 41.8 | 39.1 | 38.5 | — |
| Vision | OmniDocBench | — | 91.1 | 89.8 | 85.8 | 87.9 | 89.4 | — |
| Vision | PerceptionBench | — | 58.5 | 57.2 | 59.7 | 47.2 | 55.8 | — |
Kimi K3 combines Kimi Delta Attention and Attention Residuals with a highly sparse Stable LatentMoE architecture, activating 16 of 896 experts. Moonshot AI also uses quantization-aware training with MXFP4 weights and MXFP8 activations for broad hardware compatibility.
Choose Kimi K3 from the Cloud model picker in Bionic. LM Studio Cloud inference servers are US-based, and Zero Data Retention is enabled by default: your prompts and outputs are not retained or used for training.
Read LM Studio's Kimi K3 announcement · Read Moonshot AI's Kimi K3 technical blog