Kimi K3
The world's first open 3T-class model (2.8T MoE), engineered for autonomous long-horizon coding, CAD, and deep research.
Generic Info
- Publisher: Moonshot AI
- Release Date: 2026
- Parameters: 2.8 Trillion Total (16 active out of 896 LatentMoE experts)
- Context Window: 1,000,000 tokens (1M native)
- Architecture: Kimi Delta Attention (KDA) + Attention Residuals (AttnRes)
- License: Kimi Community Open-Weights License
- Key Capabilities: Autonomous Multi-Hour Coding, GPU Kernel Optimization, CAD/Chip Design, Vision-in-the-Loop
Kimi K3 redefines what open weights can achieve by scaling to 2.8 trillion total parameters with high MoE sparsity. Utilizing Kimi Delta Attention (KDA) and Attention Residuals, it achieves a 2.5× scaling efficiency advantage over previous generations. Operating with minimal human oversight, Kimi K3 sustains multi-hour terminal sessions, navigates massive mono-repos, and synthesizes interactive research widgets and full software architectures.
Hello World Guide
Run inference with Kimi K3 using Hugging Face transformers or compatible MoE backends.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "moonshotai/Kimi-K3"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
prompt = "Write an optimized Triton kernel for flash cross-attention with causal masking."
messages = [
{"role": "system", "content": "You are Kimi K3, an expert autonomous engineering agent."},
{"role": "user", "content": prompt}
]
inputs = tokenizer.apply_chat_template(
messages,
return_tensors="pt",
add_generation_prompt=True
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=1024,
temperature=0.7
)
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
Industry Usage
Deep Engineering Sessions
Navigates million-line codebases, orchestrates terminal commands, and resolves complex compiler errors without human intervention.
Hardware & Chip Design
Specialized domain fine-tuning enables Verilog generation, synthesis constraint verification, and CAD workflow automation.
Interactive Knowledge Dashboards
Synthesizes comprehensive multi-domain analyses complete with executable widgets, data visualizations, and formal reports.