ElevenLabs Eleven v4 and Turbo (Sept 28, 2026): 90+ languages, instant voice clones. How modern expressive TTS works and how to budget voice-agent latency.
Alibaba Qwen3.8-Omni-Flash (Sept 18, 2026) takes text, image, audio and video natively with 1M context at $0.15/$0.47 per M tokens. Architecture and cost analysis.
OpenAI Astra launched Sept 3, 2026 for computer use and cyber. How opaque recurrence works, why it weakens chain-of-thought monitoring, and what builders should do.
PrismML released Ternary Bonsai 2 27B under Apache 2.0 at 1.76 effective bits per weight. What ternary weights mean for edge memory, latency and accuracy.
Xiaomi released MiMo-V2.6 Pro and Flash under MIT license with its RL training stack open-sourced. Architecture, benchmarks, cost, and self-hosting notes.
Gemini 3.8 Flash launched Sept 2, 2026 with a 1M context window and introductory pricing that doubles on Jan 1, 2027. A cost-per-task breakdown for builders.
OpenAI shipped GPT-6 Sol and Luna on Sept 22, 2026 with a 1.05M context window and prices 50% below GPT-5.6. Here is what changed and how they compare.