AWQ
A post-training quantization method for fast, accurate GPU inference.
A post-training quantization method for fast, accurate GPU inference. (M08, M14)
A 4-bit weight-only method that protects the most important weights. Good quality-for-size; well supported by vLLM.