Local Inference

Notes on running your own LLMs at home: hardware, quantization, and the tool stack, without a hosted API in the loop.