Kimi K3

8 stories

Kimi K3 is a large language model reported to have 2.8T parameters and roughly 1.45 TB of expert weights. Recent coverage highlights commercial and engineering use: Moonshot plans to leverage Kimi K3 toward $2 billion annualized sales, Cognition post‑trained it to produce the SWE‑2 coding model, and users measured it running ~1 token/s on an M5 Max MacBook Pro with four SSDs.

Related topics

Report: China narrows AI gap with U.S.

The report says China has been rapidly closing its AI lead over the U.S. in model capabilities and research despite U.S. export restrictions on advanced chips and equipment. Experts describe the U.S.–China model gap as narrowing, noting Chinese efforts such as Moonshot AI’s Kimi K3.

ABC News - Breaking News, Latest News and Videos · · Details

Open Chinese models narrow gap with frontier models

A Mozilla report finds that the performance gap between Chinese open-weight models and US frontier models has narrowed to roughly 4.4 months. It highlights Moonshot AI's Kimi K3 scoring three points behind Anthropic's Fable 5 on a composite index while costing about 30% as much, suggesting open models are preferable for most routine workloads.

Ars Technica · · Details

CrofAI exposed as OpenRouter wrapper with big markups

Researchers allege CrofAI (NahCrofAI) was an OpenRouter wrapper that silently routed requests to cheaper/weaker models while charging up to 20x markups. The operator denied then backtracked and ultimately wiped the service; examples include routing Kimi K3 requests to GLM models, per the exposé.

r/LocalLLaMA · · Details

Do tech reports matter for PhD apps

A Reddit post asks whether tech reports for large models (not arXiv papers) — e.g. Kimi K3, DeepSeek, Gemini, Mistral Leanstral — carry similar weight to a first‑author A* paper in PhD applications, seeking community input.

r/MachineLearning · · Details

Cognition Unveils SWE-2 Coding Model

Cognition launched SWE-2, a coding model scoring 50.0 on FrontierCode 1.1 Main—within one point of Fable 5.1 while 64% cheaper; it was post-trained from the 2.8T-parameter Kimi K3 and uses single-run RL across reasoning levels to improve cost-performance.

Hacker News · · Details
That is everything