Sorry for the slop, but I was impressed by this model as I have been testing Qwen3.8-27B Q8_0 locally on my ROG Flow Z13 (Ryzen AI Max+ 395, 128 GB unified memory) and this model was the only one who could made this short simulator (and I have tested a lot of models). Prompt: "Create a beautiful, relaxing flight simulator in a single HTML page." It generated the whole thing through an agent using file/bash tools. Setup: Lemonade Server + llama.cpp ROCm Q8_0 weights + Q8 KV cache Native MTP speculative decoding 64 GB VRAM / 64 GB RAM split ~142k context I'm seeing roughly 9-19 tok/s depending on the agent step, with some generations sustaining 16-19 tok/s and MTP acceptance reaching 97-99% and it took ~20min to generate this simulator. submitted by /u/seti_at_home [link] [comments]
阅读原文 ↗ 内容来自 Reddit r/LocalLLaMA(社区)