Models

Kimi K3 redraws the model map with a 1M-token context window

Moonshot AI has moved its flagship to a 2.8-trillion-parameter, native-vision model built for long-horizon coding, knowledge work, and reasoning.

The read

K3 is now the default model to evaluate for difficult Kimi workloads. The important changes are the 1M-token context window, native vision, new attention architecture, and stronger long-running agent work.

What changed

Kimi K3 is Moonshot AI's new flagship across Kimi, Kimi Work, Kimi Code, and the API. It is a 2.8T-parameter model with native visual understanding and a context window of up to 1,048,576 tokens.

The architecture moves beyond the K2 family with Kimi Delta Attention and Attention Residuals. Moonshot says the model is designed for long-horizon software work, deep knowledge tasks, and extended reasoning rather than short chat alone.

  • Model ID: kimi-k3
  • Context window: up to 1M tokens
  • Inputs: text and vision
  • Reasoning effort: low, high, or max
  • Availability: Kimi, Kimi Work, Kimi Code, and the Kimi API

The open-weight timing matters

Moonshot announced that full K3 weights are scheduled for July 27, 2026. Until that release lands, the official hosted products and API are the authoritative way to evaluate the launch model.

That distinction matters for benchmark claims and third-party deployments. Kimi Insider will mark the weights as pending until the repository and technical report are public.

Who should test it first

Teams working with large repositories, long document sets, visual engineering tasks, or multi-hour agent runs have the clearest reason to test K3. Straightforward coding sessions may still be better served by the more focused K2.7 Code options, especially when latency and quota efficiency matter.