The weights I've tried is the "DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix" ones which fits just about within 96GB VRAM. With some tuning, I've managed to get it to get up to ~60 t/s, I'm sure there is other things to do there too :)