Kimi K3 explained: Moonshot AI's 2.8T-parameter open-weight MoE - Kimi Delta Attention, 1M-token context, training, benchmarks, deployment cost and honest limits of the reasoning model in 2026.
Physical Intelligence pi0.5 explained: the ~3.3B vision-language-action robot foundation model - architecture, FAST action tokenization, training on 10+ embodiments, benchmarks, deployment and honest limits in 2026.
Mistral Large 3 explained: the 675B/41B Apache-2.0 sparse MoE with a 262K context window - architecture, training, benchmarks, licensing, self-hosting VRAM, pricing and honest limits in 2026.
OpenAI Sora 2 explained: the diffusion-transformer video model - spacetime patches, MM-DiT with synchronized audio, latent compression, clip length, physics, deployment, pricing and honest limits in 2026.
A deep dive on Moonshot AI Kimi K2: the Mixture-of-Experts architecture, training recipe, agentic and coding benchmarks, open weights, license, and how it compares to peers.
Google Gemini 3.5 Pro explained: the 2M-token context flagship, architecture, training, benchmark scores, pricing, and how it compares to GPT-5.6 and Claude.
Google Gemini 3.5 Flash explained: the MoE multimodal architecture, context window, real 2026 benchmarks, pricing, latency, and how it compares to GPT and Claude.
DeepSeek V4 explained: the 1.6T-parameter MoE architecture, Compressed Sparse Attention, 1M-token context, SWE-bench and reasoning benchmarks, pricing, and how to deploy it.