← Back to overview

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs