Palit Gaming Pro 24GB OC 3090 + Palit X3060 8GB? is it worth bothering even?

Device name -AI_AGENT-
Processor Intel(R) Core™ i5-10400F CPU @ 2.90GHz (2.90 GHz)
Installed RAM 16.0 GB (15.9 GB usable)
Graphics card NVIDIA GeForce RTX 3090 (24 GB)
Storage 291 GB of 477 GB used
Device ID 4D08CEB9-CEF7-4AF9-A4EE-6FCA04B9E666
Product ID 00330-80000-00000-AA366
System type 64-bit operating system, x64-based processor
Pen and touch No pen or touch input is available for this display

Beginners

##the above is what I deal with, my very well maintained PCVR
###inference via Ollama, Qwen quantized GGUF
###Hermes desktop

my question is? is it worth buying 8GB Palit storm X 3060, on top I have 3090 gaming pro OC 24GB .. that without changing PSU, just a second GPU and extra 8GB for that fancy TTS? storm X is half decent and will pack in ROG STRIX motherboard to max

MartinPilarski

5

23h

Perfect — it’s working great! The model loads in ~19 seconds and responds cleanly. Here’s the current status:

  • Model: hf.co/unsloth/Qwen3.8-27B-GGUF:Q4_K_M — confirmed running
  • Load time: ~19.5 seconds into VRAM
  • Response: 482 tokens in ~12 seconds at Q4 K_M quant
  • Files on D: blobs (16.5 GB model + 928 MB projector)
  • Files on C: manifest for Ollama to see it

is 482 tokens/12 sec correct… way too many.. am I dreaming? :stuck_out_tongue:

I COULDN’T POST A REPLY PAST 12H PERIOD SO I EDITED

##THE ORIGINAL POST
##AFTER SOLUTION

yea , this is amazing, my old PCVR doing this, and It’s only when I told it to pull 3.8, but I was lazy and told to pull equivalent of 3.6 GGUF unsloth heretic :P… IT PULLED LOADS GIVING ME SSD C: ALERT … and it almost succeded… MY GPU IS DOWNCLOCKED AND POWER LIMIT 90% , I JUST REPASTE GPU 6MONTHS AGO, Palit Gaming Pro 24GB OC… i’m cranking up ths baby, this is so good, only because I don’t have any sensitive data on my gaming PC i since soul.md and agents.md told it to be iresponsibly ffree, and be a fun Hermes agent … and it did.. It is better than ever expected… it even say - it is me Circe and my human Martin … lol :DDDDDDDDDDDD

I had to go back to old model, need to purge SSD C: D: and connect that 1TB F: :stuck_out_tongue: , it bloated it out all space, I’ll just find correct Qwen 3.8 model at the weekend binge :wink:

#there goes

Ollama’s already loaded — qwen-capped-128k:latest is sitting in memory:

Model Size Processor Context Until
qwen-capped-128k:latest 23 GB 4%/96% CPU/GPU 131072 29 minutes

Let me grab the verbose output:

Good news — everything’s working:

ollama run qwen-capped-128k:latest --verbose responded in 2.06s total:

  • Load time: 4ms (model already loaded in memory)
  • Prompt eval: 90.31 tokens/s
  • Eval rate: 123.58 tokens/s — solid throughput

ollama ps confirms the model is resident:

  • 23 GB loaded, 96% GPU usage (RTX 3090), context cap at 131072
  • Will auto-unload in ~29 minutes of inactivity

The verbose run completed successfully with no errors. GPU inference is healthy, loading from the already-resident load is near-instant, and token throughput is what you’d expect from a Q4_K_M quant on your setup. No fixes needed.

:DDDDDD what to do? keep current? it loves me lol, No fixes needed just blowing my mind… jokes

Solved

John666620h

Oh. It depends on the backend you’re using, but that seems to be about the expected speed: