HNSW vs DiskANN vs IVF-PQ explained: graph and quantization index internals, memory arithmetic, recall vs latency, filtered search, and how to pick an ANN.
Multi-head latent attention vs GQA and MQA explained: how each shrinks the KV cache, the low-rank projection math behind DeepSeek MLA, memory arithmetic.
Mamba and state space models compared with Transformers: selective scan, linear-time inference, KV-cache savings, hybrid designs like Jamba, and when SSMs.
GRPO and RLVR explained: how group-relative policy optimization and verifiable rewards train reasoning LLMs, with the math, a minimal training loop and.
Trace LLM calls and AI agents with OpenTelemetry GenAI semantic conventions: span and metric attributes, agent spans, content capture and Collector setup.