Risk-Routed Heterogeneous KV Memory
An investigation of KV-cache compression as fidelity routing: exact-critical spans retain full KV precision while lower-risk background context uses low-bit quantized KV.
Research question: What memory representation should each token use?
The current implementation is an inference-time quantize/dequantize proxy, not a packed low-bit storage kernel.