Moonshot AI has introduced Kimi K3, an open-source model with 2.8 trillion parameters and a one-million-token context window, targeting complex scientific and coding workflows. The company claims performance gains, but key limitations remain
Moonshot AI has announced the release of Kimi K3, an open-source large language model with 2.8 trillion parameters, positioning it as the largest open-weight model available to date. The company describes Kimi K3 as a 3T-class system designed for extended scientific and technical workflows, including research report generation, code development, and multimodal tasks that combine text and images. The model is intended to support users in research, engineering, and knowledge-intensive domains, with a particular emphasis on handling long, complex projects that exceed the capacity of previous open models.
Kimi K3 is built on Moonshot AI's proprietary Kimi Delta Attention and Attention Residuals architecture, using a Mixture-of-Experts (MoE) framework that activates 16 out of 896 experts per inference. According to the company, these architectural choices improve scaling efficiency by approximately 2.5 times compared to the previous Kimi K2 model. The model features a one-million-token context window, which is among the largest available in open models, allowing it to process and generate outputs across extended documents, large codebases, and workflows that require persistent memory over long sessions.
Evaluation and Demonstrations
Moonshot AI reports that Kimi K3 outperformed several leading models in its internal benchmark suite, including OpenAI's GPT-5.5, Anthropic's Claude Opus 4.8, and Zhipu AI's GLM-5.2. The company demonstrated Kimi K3's capabilities on tasks such as optimizing GPU kernels, developing a GPU compiler (MiniTriton), designing a prototype AI chip using open-source electronic design automation tools, and reproducing computational astrophysics workflows. In one demonstration, Kimi K3 reportedly completed a research task in about two hours that would typically require one to two weeks of work by an experienced researcher. However, Moonshot AI acknowledges that Kimi K3 still lags behind the most advanced proprietary models, specifically Claude Fable 5 and GPT 5.6 Sol, based on its own evaluations.
The company has not yet released independent benchmark results or peer-reviewed technical documentation. All reported performance figures and demonstrations originate from Moonshot AI's own blog and technical materials. The full model weights are scheduled for release by July 27, 2026, along with a technical report detailing the model's architecture, training process, and evaluation methodology. Until then, external verification of Kimi K3's capabilities remains limited to company-provided evidence.
Technical and Social Context
Kimi K3 is available for use through Moonshot AI's platforms, including Kimi.com, Kimi Work, Kimi Code, and the Kimi API. The model launches with maximum reasoning effort enabled by default, with plans to introduce additional effort modes in future updates. The system is positioned as a tool for scientific research, technical analysis, and knowledge work, with features such as Widgets and Dashboards that allow users to create persistent, interactive workspaces. These features are intended to support workflows that require iterative analysis, visualization, and report generation.
The release of Kimi K3 comes amid intensified competition among Chinese AI developers to close the gap with leading U.S. companies in large language model development, as reported by the South China Morning Post. The global market for artificial intelligence is projected to expand rapidly, with significant implications for labor markets and research practices. However, the practical impact of Kimi K3 will depend on independent evaluation, real-world deployment, and the extent to which its open-source status enables broader scrutiny and adaptation.
Numerical Context
Kimi K3's 2.8 trillion parameter count makes it the largest open-weight model announced to date, according to Moonshot AI. The model's one-million-token context window is designed to support extended workflows, such as reviewing more than 20 scientific papers, evaluating over 300 equations of state, and generating thousands of lines of code in a single session. The Mixture-of-Experts architecture activates 16 out of 896 experts per inference, which the company claims improves scaling efficiency by 2.5 times over Kimi K2. All performance comparisons and efficiency claims are based on internal company benchmarks and have not yet been independently verified.
Understanding model parameters is essential for interpreting claims about large language models. A parameter in this context refers to a learned weight in the neural network that determines how the model processes and generates text or other data. While higher parameter counts can enable more complex pattern recognition and longer context windows, they do not guarantee improved accuracy, reliability, or safety. Model performance depends on architecture, training data, evaluation methods, and real-world deployment conditions. Open-weight releases allow external researchers to audit, adapt, and test models, but independent verification is necessary to establish practical capability and risk.