LOGBOOK

Kimi K3: what does an open 2.8-trillion-parameter AI model mean?

Moonshot AI presents Kimi K3 as its most capable model yet. We examine its enormous scale, context window, development potential, and the important caveats.

Kimi K3: what does an open 2.8-trillion-parameter AI model mean?

On July 16, 2026, Moonshot AI introduced Kimi K3. The company calls it its most capable model yet and the first open model in the three-trillion-parameter class. That is a striking step, but parameter count only tells part of the story. More interesting is what this scale may enable for long-running development tasks, large amounts of context, and multimodal work.

A 2.8-trillion-parameter model

According to Moonshot, Kimi K3 has 2.8 trillion parameters. It is a Mixture of Experts model, which means the complete network is not used for every step. Moonshot describes an architecture that activates 16 out of 896 experts. This allows the model to be extremely large without using all of its capacity for every response.

The architecture combines Kimi Delta Attention with Attention Residuals. Moonshot claims this approach delivers roughly 2.5 times better scaling efficiency than Kimi K2. This is a vendor claim that still needs broader independent evaluation.

Native vision and up to one million tokens of context

K3 processes text and images and, according to the official documentation, supports up to one million tokens of context. In theory, this allows it to work with large codebases, long case files, and extensive project histories in a single session. In practice, the available context depends on the product and subscription. The official Kimi Code documentation, for example, states that not every membership tier includes the full one-million-token window.

A large context window is also no guarantee that every detail will be understood equally well. Reliable product work still requires a clear brief, good source selection, and human review.

Built for coding, agents, and knowledge work

Moonshot positions K3 for long-horizon programming, complex engineering, reasoning, and knowledge work. The model is available through Kimi.com, Kimi Work, Kimi Code, and the API. Within Kimi, it can also produce documents, presentations, and spreadsheets and perform agent tasks.

For product teams, the combination is especially interesting: a model that can retain a large amount of project context, understand visual information, and use tools across multiple steps. This could support research, analysis, prototyping, code review, and documentation.

How does K3 compare with Fable 5 and GPT-5.6 Sol?

The honest answer is that there is no single winner. On the Artificial Analysis Intelligence Index published on July 17, Kimi K3 scored 57, compared with 60 for Claude Fable 5 and 59 for GPT-5.6 Sol. On DeepSWE v1.1, which contains 113 long-horizon engineering tasks, K3 scored 69% plus or minus 5 percentage points. Fable 5 scored 70% plus or minus 4 and Sol scored 73% plus or minus 3. The uncertainty ranges overlap, so small ranking differences should not be overstated.

K3 reversed the order in visual web development. On July 16, it provisionally led Arena WebDev with 1679 points, ahead of Fable 5 at 1631 and Sol at 1618. That leaderboard is based on human preference across real web tasks, but K3's result was still marked as preliminary.

The type of work matters too. On AA-Briefcase, a private benchmark for long-horizon knowledge work, K3 reached an Elo of 1543. That placed it behind Fable 5 at 1574 but ahead of Sol at 1501. The picture is nuanced: K3 is not universally better than Fable or Sol, but it demonstrably belongs in the same frontier group and wins in several relevant categories.

Why this release matters

The main significance of Kimi K3 is not that every company should now run a 2.8-trillion-parameter model. A model at this scale requires enormous infrastructure. Its importance lies in the pressure a capable open model puts on the market. Better access to model weights and architectural knowledge can accelerate research, reduce vendor dependence, and give organisations more choice over where their data is processed.

That freedom can be relevant for European organisations. An open model creates options for private hosting or specialised deployments, but only when the infrastructure, licence, security, and total cost demonstrably fit the use case.

A Chinese model calls for nuance, not reflexes

Moonshot AI is based in Beijing. That matters when an organisation uses Kimi as a cloud service, especially with sensitive information or in a regulated sector. Review the contracting entity, legal jurisdiction, data location, access controls, retention periods, and audit options. This due diligence belongs with every non-European AI provider, but geopolitical tensions and Chinese regulation may create additional continuity and compliance risks.

At the same time, origin alone is not a judgement about model quality. Strong Chinese competition benefits the market: it puts pressure on prices, accelerates open research, and prevents a small number of American companies from controlling the entire AI infrastructure. If the announced weights become genuinely available and usable, private hosting could reduce some data risks. It would not remove questions about model provenance, training data, licensing, or infrastructure.

Important caveats

  • Most performance figures currently come from Moonshot and do not yet amount to a fully independent verdict.
  • Moonshot announced that the full model weights would become available on July 27, 2026. At this article's publication date, July 22, that date is still in the future.
  • The one-million-token context window is not available in every Kimi product or membership tier.
  • Self-hosting a model at this scale requires costly and specialised infrastructure.
  • Even a strong model can make mistakes. Human direction, testing, and security reviews remain essential.

Our conclusion

Kimi K3 shows that open AI models are moving closer to the frontier occupied by leading commercial systems. Its combination of scale, native vision, long context, and agent capabilities is technically impressive. Its real value will become clearer through independent testing and through products where these capabilities demonstrably save time or improve quality.

For organisations, the right question is therefore not which model has the biggest numbers. It is which model, with the right human direction, safely and affordably produces the best result for the actual product.

Sources

All hands on deck?

Tell us what you're building. We'll tell you what we can ship by Friday.