An interactive, step-by-step diagnostic guide to troubleshoot slow token generation, high prompt evaluation latency, and CPU bottlenecks when running local LLMs in Ollama or LM Studio. Follow guided workflows to optimize GPU layer offloading, adjust context sizes, and select proper model quantization levels.