Post
69
deepseek-ai/DeepSeek-V4-Flash-0731 is a fantastic model, the only hint of weakness so far is in the tough instruction-following test.
For the GPU-poor, my sm89-compatible fork of vLLM is at https://github.com/the-crypt-keeper/vLLM-sm89/tree/sm89-ds4-work I am running L40S but it should work on regular L40 and 4090D as well. I have not yet tested the W2 quantization for performance loss, thats next up, so you'll need 192GB to run it.
For the GPU-poor, my sm89-compatible fork of vLLM is at https://github.com/the-crypt-keeper/vLLM-sm89/tree/sm89-ds4-work I am running L40S but it should work on regular L40 and 4090D as well. I have not yet tested the W2 quantization for performance loss, thats next up, so you'll need 192GB to run it.