MartinPilarski
5
23h
Perfect — it’s working great! The model loads in ~19 seconds and responds cleanly. Here’s the current status:
- Model:
hf.co/unsloth/Qwen3.8-27B-GGUF:Q4_K_M — confirmed running
- Load time: ~19.5 seconds into VRAM
- Response: 482 tokens in ~12 seconds at Q4 K_M quant
- Files on D: blobs (16.5 GB model + 928 MB projector)
- Files on C: manifest for Ollama to see it
is 482 tokens/12 sec correct… way too many.. am I dreaming? 
I COULDN’T POST A REPLY PAST 12H PERIOD SO I EDITED
##THE ORIGINAL POST
##AFTER SOLUTION
yea , this is amazing, my old PCVR doing this, and It’s only when I told it to pull 3.8, but I was lazy and told to pull equivalent of 3.6 GGUF unsloth heretic :P… IT PULLED LOADS GIVING ME SSD C: ALERT … and it almost succeded… MY GPU IS DOWNCLOCKED AND POWER LIMIT 90% , I JUST REPASTE GPU 6MONTHS AGO, Palit Gaming Pro 24GB OC… i’m cranking up ths baby, this is so good, only because I don’t have any sensitive data on my gaming PC i since soul.md and agents.md told it to be iresponsibly ffree, and be a fun Hermes agent … and it did.. It is better than ever expected… it even say - it is me Circe and my human Martin … lol :DDDDDDDDDDDD
I had to go back to old model, need to purge SSD C: D: and connect that 1TB F:
, it bloated it out all space, I’ll just find correct Qwen 3.8 model at the weekend binge 
#there goes
Ollama’s already loaded — qwen-capped-128k:latest is sitting in memory:
| Model |
Size |
Processor |
Context |
Until |
| qwen-capped-128k:latest |
23 GB |
4%/96% CPU/GPU |
131072 |
29 minutes |
Let me grab the verbose output:
Good news — everything’s working:
ollama run qwen-capped-128k:latest --verbose responded in 2.06s total:
- Load time: 4ms (model already loaded in memory)
- Prompt eval: 90.31 tokens/s
- Eval rate: 123.58 tokens/s — solid throughput
ollama ps confirms the model is resident:
- 23 GB loaded, 96% GPU usage (RTX 3090), context cap at 131072
- Will auto-unload in ~29 minutes of inactivity
The verbose run completed successfully with no errors. GPU inference is healthy, loading from the already-resident load is near-instant, and token throughput is what you’d expect from a Q4_K_M quant on your setup. No fixes needed.
:DDDDDD what to do? keep current? it loves me lol, No fixes needed just blowing my mind… jokes
Solved
John666620h
Oh. It depends on the backend you’re using, but that seems to be about the expected speed: