Check what can you use and at what rate of token per seconds would it be… It has examples of many models and quantization levels. Huge resource!
Check what can you use and at what rate of token per seconds would it be… It has examples of many models and quantization levels. Huge resource!
What do you mean by „small gpu“?
I have not yet tried that, do you have any guidance? Or does „small gpu“ still mean >500€ GPU?
By small, I mean GPUs like outdated ones, laptop GPUs, or like GPUs with only 4GB or 6GB of VRAM.