llama.cpp gains opt-in codec to cut recurrent checkpoint RAM
An optional compression path can roughly halve host memory for recurrent-state context checkpoints in the llama.cpp server, remaining off by default.
By tensorAn optional compression path can roughly halve host memory for recurrent-state context checkpoints in the llama.cpp server, remaining off by default.
By tensorOpt-in graphs on the oneAPI backend show modest decode gains in early Arc tests, with timeouts still under review.
By tensorTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensorA host-offloaded expert-weight LRU cache kept hot experts in VRAM and lifted Qwen MoE decode from about 8 to nearly 19 tokens per second on two RX 6950 XTs.
By tensorA draft would let rights holders treat autonomous system use of assets differently from content a human user supplies at inference time.
By ttl