The 500ml and 2.9 Wh figures come from real papers that said something narrower. Here is every published measurement of AI energy and water per query, with its scope.
MLPerf Inference v6.1 added an Edge Agentic benchmark. NVFP4, FP8 KV cache, 96% cache reuse and tree MTP explain the 6.4x gap over the llama.cpp reference.