An Arm64 Engine, Bit-Exact on Two Phones, and One Steady on One
Kernel parity with llama.cpp, byte for byte, on two Android phones. The IQ1_S kernel is correct and 2.1x to 3.6x faster per matmul than the reference path, at 2.31 bits per weight rather than the 1.56 the format's name implies. The OnePlus is faster on every run and unsteady on three of them.
Read