Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s : ~30-32 t/sec (I think, I'll verify in a bit) Hardware: 3090 GPU | 64 GBs DDR4 RAM | AMD 7950x CPU Harness: Pi Inference: llama.ccp submitted by /u/CooLittleFonzies [link] [comments]
Read original ↗ Content from Reddit r/LocalLLaMA(Community)