My experience with heavily quantized models is they do better than a similar sized unquantized model but don't really perform anywhere near the pre-quantized model.
>My experience with heavily quantized models is they do better than a similar sized unquantized model but don't really perform anywhere near the pre-quantized model.
People have tested it. Q8 has essentially no drop in quality, Q4 is measurable but still not realistically a problem. If this impacts you, just pay for the commercial saas option.
This assumes that the benchmarks are representative of real usage scenarios. I'm not saying that there is bad faith, but that benchmarking is really hard in the context of LLMs.
>This assumes that the benchmarks are representative of real usage scenarios. I'm not saying that there is bad faith, but that benchmarking is really hard in the context of LLMs.
It's a fair point, but the conclusion is 'i dont know'
I could assume that it gets better because it'll keep to simpler code.
So in around 6 months, we will see that the person who bought this H200 in the listing just got scammed for $250k and will realize that you just needed specific quantizations to the model and a few optimizations to run locally.
Unless they want to train their own model, buying this for inference for $250k is unnecessary and still isn't enough for a full production deployment.