[Return]

Report a post

Preview
>>25129
12gb is pretty small memory even if you ignore the model size itself there is not much left for context memory

You probably should use 7b models based on llama2 and you may try llama.cpp combining CPU + GPU
those small models are pretty fast even when running on CPU
Post number No.25130
Board Literature
Optional. Describe what's wrong with it.