Magnitude auto-tunes local AI inference for your exact chip, claims 2x over llama.cpp — type0 | type0