Yes! your gaming gpu runs an autonomous local ai agent, here is what to run on 8, 12, 16 and 24gb
two thirds of gaming pcs on steam have 8 to 16gb of vram, and an rtx card in that range now runs a 27b, 42 tok/s on 8gb, 50 tok/s on 12gb with vision and 67 tok/s on 16gb with the full 262k window.
i did not buy a new card. a lab compressed a 27b down to 5.95 gb, i fixed the kernel that was holding it back on gaming cards and grafted its speed head back on, then i ran it on every vram tier i could get my hands on. this is what i found, card by card, with the settings and the catches, and every number in it comes from a run i can show you.
i packaged the model so you can run it too, with its speed head plus prebuilt llama.cpp for rtx 30, 40 and 50 cards, and it has 27k downloads so far:
start with the card you actually own
steam's hardware survey is the closest thing we have to a census of gaming pcs. in september, 16gb was the most common vram tier at 27.21% after it overtook 8gb in july, 8gb sat right behind at 26.71%, 12gb at 13.06%, and 24gb and up at 7.41%. that puts 67% of steam gpus between 8 and 16gb. the single most owned card is now the rtx 5070 12gb at 5.86%, the rtx 3060 12gb that led all summer is fifth at 3.54%, and seven of the top fifteen cards have 8gb.
if you are reading this on a gaming pc, you are probably in that 67%, and that is who this guide is for. and it is enough to make you realize what it feels like to really own your thinking end to end, from harness to model weights to hardware, and it will be a single shift in your mind that will trigger wanting to add one more gpu, and the quest begins. when it does, let it. it is the right call.
the model, and the two fixes that made it fast
for 8 to 16gb i tested one model for all of us, the same file on every tier so the numbers compare like for like, and the map after this section shows everything else that fits your card, text and image.
bonsai 2 27b is prismml's ternary compression of qwen 3.8 27b, the same model i crowned on the 3090 in august, squeezed into 5.95 gb. their own evaluation puts it at 98.2% of the full model, it keeps qwen's 262k window and its vision, and it is the #1 trending text model on hugging face right now with 3.97 million downloads. one catch up front: it needs prismml's fork of llama.cpp, because stock llama.cpp loads the file and prints gibberish.
out of the box on a 3060 it decoded about 25 to 26 tok/s. the fork's kernel for this format was leaving most of the card idle, so i fixed the kernel and the same file went to 40 tok/s. then i grafted qwen 3.8's own mtp head back onto it, the small draft head that guesses the next token so the big model checks two in one pass, and it hit 50 tok/s.
here is the whole story on one card, the same file at five depths, stock in gray, my kernel in teal, the kernel plus the head in orange, from a fresh chat out to 119k tokens of context:
with batch invariance on, the text comes out byte-identical to running it without the head, so the speed costs you nothing in quality. the kernel is pull request #218 on prismml's fork, approved and waiting to merge, and the merged model file is on hugging face with 27k downloads.
the night it shipped:
Übersetzt (Originalsprache Englisch)
deine RTX 3060 hat heute Morgen Bonsai 2 mit 26 Tok/s laufen lassen und jetzt macht sie 40 Tok/s, ich habe den Tag in der PrismML Llama.cpp Fork verbracht, damit deine RTX 3060 jubeln kann.
der Kernel, der die Weights liest, hat zwei Drittel der Threads der Karte bei den größten
Zitat
Sudo su
@sudoingX
9:49
Übersetzt (Originalsprache Englisch)
Die Timeline hat zwei Tage lang gesagt, Bonsai 2 könne nicht bauen. Hier sind 1 Stunde 24 Minuten davon, wie es baut, in einem Take, aus einem Absatz, auf einer RTX 3060 12GB, auf 8x beschleunigt, damit ihr das Ganze anschauen könnt.
Was ihr anschaut, ist eine 5,9 GB ternäre x.com/sudoingX/statu…
here is the map before the tiers, everything that fits each card, every tier runs everything to its left, and a star marks the ones i have run:
8gb: yes, it is an agent card now
cards: rtx 3050, 3060 ti, 3070, 4060, 5060, the 8gb 4060 ti and 5060 ti, and the 4060 and 5060 laptops.
run: bonsai 2 27b, mtp head off, a 96k window.
the 8gb answer came from a contributor, not from me. they ran it on their own rtx 3060 ti and sent the numbers to my repo as a pull request: the full 262k window loads in 7.7 gb, and it holds 42.8 tok/s all the way out to a 96k window, then falls off a cliff past about 112k. so 96k is your line, and dropping to 64k buys you almost nothing, unless the card also drives your monitor, then take 64k for the headroom. vision fits too, at a 32k window, and you do not need a second download, the same file with the head switched off uses the same vram as prismml's plain one.
here is that cliff on the 3060 ti, flat at 42.8 tok/s from 64k to 96k, then both builds collapse at 112k:
before bonsai 2, my 8gb pick was gemma 4 e4b. in may i ran three models on a gtx 1080 8gb through hermes agent, and gemma won on speed and window, 42.13 tok/s with a 656k window thanks to its sliding window attention, with qwen 3.5 9b second at about 30 tok/s and 248k. gemma is still the pick if you want a small fast model with a giant window. bonsai 2 is the pick if you want a 27b brain on the same card.
12gb: the rtx 3060 runs a 27b at 50 tok/s
cards: rtx 3060, 4070, 5070, 3080 12gb, 4070 super, 4070 ti.
run: bonsai 2 27b with the kernel and the mtp head, a 131k window, vision on top, 11.5 gb.
this is the tier i lived on for a week. with the head on, a fresh chat runs at 50.1 tok/s, a photo prompt at 50.05 tok/s with the vision tower loaded at a 131k window, and 40 tok/s with the head off. the head keeps paying as the context fills, 61% faster than stock at 18k tokens in and still 32% faster at 119k. it drives a browser too, 8 of 10 web tasks done, 29% faster than the stock build.
then i gave it the octopus invaders prompt, the same one the 9b got in march, through hermes agent. five hours, 328k tokens written, eight js files, zero handwritten code, 50 tok/s fresh and 22 tok/s on average across a session that ran 125k tokens deep.
Übersetzt (Originalsprache Englisch)
das ist, was 12gb vram im jahr 2026 leistet, pure zauberei
> rtx 3060 12gb, #1 gpu auf steam
> bonsai 2 27b + mtp, 5,95 gb gewichte
> hermes agent, 5 stunden, 328k token geschrieben
> 8 js-dateien, 2.368 zeilen, null handgeschriebener code
> 50 tok/s frisch, 22 tok/s im
Zitat
Sudo su
@sudoingX
Übersetzt (Originalsprache Englisch)
Verdammt, du wirst nicht glauben, was bonsai2 mit nur 12 GB VRAM draufhat, Ergebnisse kommen gleich
other 12gb owners are already past me. a 4070 12gb hit 103 tok/s with the head at a 262k window on prismml's integration branch, which carries this kernel.
16gb: the whole thing on one card
cards: rtx 5060 ti 16gb, 4060 ti 16gb, 5070 ti, 4070 ti super, 4080, 5080.
run: bonsai 2 27b with everything on, the full 262k window, the mtp head and vision, in 15.2 of 15.9 gb.
16gb is where it stops being a compromise. on a 5060 ti the full 262k window, the mtp head and the vision tower all load together, and a fresh chat runs at 67.3 tok/s with the head, 53.4 without, against 42.0 on prismml's stock build.
the window holds all the way to the end. at 261,000 tokens of context it still decodes 12.07 tok/s with the head and 12.35 without, while the stock build crawls at 3.73, more than three times slower. and it serves more than one of you, four requests at once add up to 124.7 tok/s with the head, and four agents, each with its own 60k context, run side by side at 42 tok/s total, about 10 to 11 each. all of it at about 145 watts while it decodes, roughly 0.6 kwh per million tokens.
the whole 16gb card on one chart, fresh speed, the full window, four at once, and what fits:
the night it landed on the card:
every single rtx 5060 ti 16gb vram owner, bonsai2 27b dense flies on your gpu with the full 262k context window, the mtp head and vision all loaded at once, and it runs at 67 tok/s.
> 262k context + mtp head + vision on one card: 15.2 of 15.9 gb
> 67 tok/s with the mtp head, 71
Zitat
Sudo su
@sudoingX
if you own a rtx 5060 ti or any 16gb gpu card, read this.
your card can now hold an agentic 27b with the mtp head, the vision tower and up to the full 256k context window, all at once, on one 16gb gpu. i have the numbers in hand and i'm rechecking every one against the raw logs x.com/sudoingX/statu…
and what it builds. one prompt, one session, 27,604 thinking tokens at 53 tok/s, 10,151 tokens written at 47 tok/s, 87% of the head's guesses accepted:
watch anon! right now in 2026 rtx 5060 ti 16gb vram runs 27b ai model that builds gpu monitoring ui as instructed in one shot at 67 tok/s fresh.
and i ask you what more do you want from $600, a step to own your thinking. a step to not let openai or anthropic control what you
Zitat
Sudo su
@sudoingX
0:55
every single rtx 5060 ti 16gb vram owner, bonsai2 27b dense flies on your gpu with the full 262k context window, the mtp head and vision all loaded at once, and it runs at 67 tok/s.
> 262k context + mtp head + vision on one card: 15.2 of 15.9 gb
> 67 tok/s with the mtp head, 71 x.com/sudoingX/statu…
a whole snake game in one html file, prompt to playable in 9 minutes 26 seconds, and one bug report fixed in one reply:
Übersetzt (Originalsprache Englisch)
Ich hätte das von bonsai2 auf meiner rtx 5060 ti 16gb nicht erwartet, anon, es hat ein vollständiges Snake-Spiel gebaut, das Spiel hat nicht gestartet, und es hat seinen eigenen Fehler in einem Versuch behoben.
Die erste Antwort war das Ganze in einer einzigen HTML-Datei,
Zitat
Sudo su
@sudoingX
2:53
watch anon! right now in 2026 rtx 5060 ti 16gb vram runs 27b ai model that builds gpu monitoring ui as instructed in one shot at 67 tok/s fresh.
and i ask you what more do you want from $600, a step to own your thinking. a step to not let openai or anthropic control what you x.com/sudoingX/statu…
and yes, claude code runs on it fully local, its first turn of 24,753 tokens read at 880 tok/s:
Übersetzt (Originalsprache Englisch)
Du kannst Claude Code kostenlos auf einer RTX 5060 Ti 16GB laufen lassen – und das gerade jetzt, ohne API-Rechnung und mit unbegrenzter Nutzung.
Jeder Gamer mit dieser GPU auf dem Schreibtisch sitzt auf einem Coding-Agenten, der nie an eine Nutzungsgrenze stößt. Die erste
Zitat
Sudo su
@sudoingX
2:53
watch anon! right now in 2026 rtx 5060 ti 16gb vram runs 27b ai model that builds gpu monitoring ui as instructed in one shot at 67 tok/s fresh.
and i ask you what more do you want from $600, a step to own your thinking. a step to not let openai or anthropic control what you x.com/sudoingX/statu…
from a clean card to the first token took 192 seconds and four commands.
if you would rather stay on stock llama.cpp, the same qwen 3.8 27b also comes as ordinary low-bit quants that fit this tier, unsloth's ud-iq3_s at 12.0 gb and ud-iq4_xs at 14.3 gb, and they run on any 16gb card today. bonsai 2 against those on one card, same prompts, is my next run, and the result lands in this section.
24gb: run the king in full
cards: rtx 3090, 3090 ti, 4090.
run: qwen 3.8 27b at q4_k_m, the full model, with the whole 262k window in 22.2 gb.
at 24gb you do not need the compression. qwen 3.8 27b is still the king here, and at q4_k_m it fits with its whole 262k window in 22.2 gb, and on my 3090 it runs 31.0 tok/s, 41.3 with the mtp head. owners in my qwen38-mtp repo have pushed 3090s past 60 tok/s with the head on newer llama.cpp builds, so rebuild llama.cpp before you tune anything.
and leave the power slider alone. a contributor's 3090 ran 51 tok/s with the head at 250 watts and 62 tok/s at 275 watts, one slider, 22% faster for 25 watts:
Übersetzt (Originalsprache Englisch)
eine 3090 auf 250w limitiert macht 51 tok/s mit der flagge. dieselbe karte bei 275w macht 62 tok/s. ein slider, +22%, für 25 watt.
derselbe beitragende hat an dem tag zwei quants auf dieser karte getestet, Q4_K_M gegen Dynamic 3.0: 52,8 tok/s vs 52,9 tok/s, ein unentschieden.
if you are buying, a used 3090 is still the best 24gb deal, $900 to $1,200 when you look locally on facebook marketplace instead of resellers.
Übersetzt (Originalsprache Englisch)
irgendwo da draußen gibt es einen Gamer mit einer 3090, der keine Ahnung hat, dass dieselbe Karte qwen 3.8 27b dense ausführt, ein Modell, das mit bezahlten APIs mitficht.
es gibt zwei Arten von GPU-Besitzern: die, die die Karte ausnutzen, und die, die damit nur Pixel schieben.
bonsai 2 still has a job here, four agents at once or a giant window with room to spare, but for one person on one card the full model is the pick. on prismml's own card the q4 of the same model edges their compression, 98.7% against 98.2% of the full model.
a few flags and one trap
these carry across every model in this guide, not only bonsai. run it on llama.cpp and pull a fresh build before you tune anything, put every layer on the card with -ngl 99 and keep flash attention on. bonsai is the one exception on the engine, it needs prismml's fork or the prebuilt bundles on my hugging face repo, driver only, no compiler, because stock llama.cpp loads the file and prints gibberish. after the weights, the window is what eats your vram, and quantizing the k/v cache is how you buy it back, q8_0 barely costs anything and q4_0 is how bonsai fits its full 262k on a 16gb card.
keep one slot per agent and raise it only when you actually run several at once, with bonsai -np 4 on a 16gb card gives four agents 64k each. if your model has an mtp head, run it at n-max 1, n-max 2 loses to the head being off once the context fills. and check the reasoning effort your chat template ships with, qwen 3.8's template, and bonsai 2's with it, defaults to xhigh, which spends the whole output budget thinking and hands back an empty answer, so set it to medium.
that last one is not a theory, here are three build tasks at xhigh:
the long version of these settings, why each flag is there, is the post 574 of you bookmarked:
Übersetzt (Originalsprache Englisch)
Wenn du dich mit lokalem KI beschäftigst, ist der erste Fehler, den einfachen Weg zu gehen. Ollama und LM Studio sind gute Apps, aber sie verbergen genau die Teile, die du lernen solltest, sodass du am Ende Monate lang Modelle laufen lässt, ohne zu wissen, warum sie schnell,
Zitat
Sudo su
@sudoingX
Übersetzt (Originalsprache Englisch)
hier ist die Reihenfolge, in der ich in lokale KI einsteigen würde, anon:
> 12gb: lerne den Stack auf einem komprimierten 27b, wie ein Modell serviert wird, wie ein Agent über Nacht mit Features wie hermes agent's /goal läuft. du wirst gute Ergebnisse sehen und etwas Licht am x.com/sudoingX/statu…
the catches
6gb is where it breaks, only 56 of the 64 layers fit on the card and it crawls at 4.8 tok/s, fine for a chat and no good for an agent, which is why this guide starts at 8gb. the speed head runs out deep in the window, on the 5060 ti it adds 11.7% at 18k tokens, nothing by 119k and it costs 2.3% at 261k, so switch it off for giant-context jobs. and every number here comes from one card at a time, mine or a contributor's, measured between may and september on the llama.cpp build of its day, so a newer build on your card can land higher.
if you take one thing from me
you do not need a new gpu to run a real model in 2026. the card already in your pc runs a 27b today, and the model, the bundles, the prompt and the scripts are all free. pull it, run it on your card and open a pull request with your row, so the next person with your card knows exactly what to expect. and when you outgrow your card, here is the order i would climb:
Übersetzt (Originalsprache Englisch)
hier ist die Reihenfolge, in der ich in lokale KI einsteigen würde, anon:
> 12gb: lerne den Stack auf einem komprimierten 27b, wie ein Modell serviert wird, wie ein Agent über Nacht mit Features wie hermes agent's /goal läuft. du wirst gute Ergebnisse sehen und etwas Licht am
Zitat
Sudo su
@sudoingX
Übersetzt (Originalsprache Englisch)
Wenn du merkst, dass deine 12-GB-Gaming-GPU ein 27B-AI-Modell mit 50 Tok/s läuft und über Nacht agentische Aufgaben erledigt
if you get stuck anywhere, leave a comment and i will help you.
- the scripts and every card's sweep:
- the 24gb repo:
- the octopus invaders prompt:
- the kernel:
Möchtest du deine eigenen Artikel veröffentlichen?
Upgrade auf Premium
