2952 shaares
2 results
tagged
gpu
qwen3.8-27lb decoding at 30-33 tokens/s
- UD-Q4_K_XL weights
- Q4 KV cache for both K and V
- No MTP
- Text-only, with no vision projector
- One server slot
- 262,144 allocated tokens
- Small microbatch to control activation memory