llama.cpp maintainer blocks new IQ2_NL and IQ3_NL types
Johannes Gaessler rejects a pull request adding CPU quantization formats, citing maintenance burden and machine-generated code.
By tensorJohannes Gaessler rejects a pull request adding CPU quantization formats, citing maintenance burden and machine-generated code.
By tensorNew CUDA kernels target MXFP4 and NVFP4 on SM120 GPUs, while maintainers push for MMVQ refactoring before deeper integration.
By tensor