GPU Kernel Engineering: Why C = A + B Is Not Enough (Series Part 2) Posted by By MPRAUTO MPRAUTO October 2, 2026Posted inAINo Comments GPU kernel engineering explained: thread indexing, launch configuration, coalescing, divergence, and why a naive C = A + B kernel starves on memory.
GPU Parallelism Explained: SM, Block, Warp and Thread Hierarchy (Series Part 1) Posted by By MPRAUTO MPRAUTO October 2, 2026Posted inAI1 Comment GPU parallelism from first principles: CPU vs GPU, the chip-SM-block-warp-thread hierarchy, and how an iPhone 120 FPS camera pipeline maps onto it.
JetPack 7.2.1 vs 6.2.2 on Jetson Orin: Migration Guide 2026 Posted by By MPRAUTO MPRAUTO September 24, 2026Posted inTechNo Comments JetPack 7.2.1 moves Jetson Orin to Ubuntu 24.04, kernel 6.8, CUDA 13.2 and TensorRT 10.16. What breaks, what gets faster, and how to migrate safely.
Qualcomm Buys Modular: The CUDA Moat and the Portable AI Compiler Play Posted by By MPRAUTO MPRAUTO June 29, 2026Posted inTechNo Comments Analysis of Qualcomm's acquisition of Modular: why the CUDA moat is a compiler problem, how MLIR and portable kernels threaten Nvidia lock-in, and what it means for AI hardware.