They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.
(It’s not super complex code, just some scripts that I would not have taken to time to write manually.)
I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬
I meant your LLM stack lol. I just have an RX 6800 XT in my main Linux PC for inference, but it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.
What are you using? vLLM? llama.cpp? Which params? How much CPU offloading? Do you use draft models? Is it a MoE model? Have you tried llama-swap? Which agentic front-end are you using? I presume you set it up to access it without SSH’ing into the machine, did you do anything special or is it just a raw unsecured open port on the machine to the LAN?
Ohhh lol. Yeah it’s Llama-swap, running llama.cpp for now, but might add vLLM to the llama-swap config to experiment with NVFP4.
I mainly use MoE models so I can get decent speed while using a 150-200k context window. My go-to model has been Qwen3.6 35B-A3B for a while. I tried Qwen3.8 27B, but it was too slow.
Gemma4 26B-A4B also runs nice and fast, but I generally get better results from Qwen3.6. I don’t remember exactly how much CPU offloading is happening, but it’s not much. As long as I can get like 40-50 tokens/sec, I’m usually satisfied enough.
For the coding harness, I’ve been running Pi in an Apple Container (sort of like Podman, but better isolation in a microvm). Though, I recently configured VS Code to use LLMs on my server, and it was actually pretty decent. Still need to explore a bit more, but so far VS Code’s AI capabilities seem much better than they were a year ago (they seemed way behind, back then).
Also, I don’t connect any harness directly to llama-swap. I have another container running Caddy, which acts as a gateway to AI providers. For other services (e.g. OpenRouter), the API key is injected in the Caddy container. I don’t like having API keys or secrets anywhere where LLMs can read them. It’s not so bad for my own self-hosted LLMs, but not cool to send secrets to a server owned by someone else.
How has tool use been for you? I struggled a lot with tool use with Gemma and Qwen, to the point where I needed to build a healing layer.
Regarding the coding harness, I was looking for something CLI-based or JetBrains-based, and I haven’t had much luck getting my local llama.cpp models playing ball with OpenCode. They keep losing context and misusing tools.
I’m not too familiar with Apple containers as I’m running a full Linux stack, but I’ll give Pi a try, seems interesting! Does it work for coding tasks or is it strictly an “orchestrator”?
Tool use with Gemma has been hit or miss. I wouldn’t rely on it for anything unsupervised.
Tool use for Qwen3.6 has been great lately, but I do remember seeing some issues with it too, a while back. I don’t remember when/why the issues cleared up (I have tweaked configs a bit over time), but switching to Pi definitely helped.
I do remember having a lot more problems in OpenCode and it was practically unusable (which is why my recent experience with VS Code was surprising). I’d definitely recommend trying Pi.
A fresh Pi install is very minimal by design. The system prompt is tiny, so it’s a pretty good fit for small LLMs like these. It’s sort of like Neovim: Nothing fancy out of the box, but you can add lots of fancy things to it. I containerize it because I don’t like giving LLMs (especially these small ones) unrestricted access to my host computer – though, I have not seen any signs of it accidentally doing something destructive, which is surprising.
There are similar alternatives to Apple Container for Linux (e.g. Docker Sandboxes, muvm, Firecracker). There’s also this thing made specifically for Pi called Gondolin. I haven’t tried it yet, but I may end up switching to that if it could simplify my stack.
Here’s my current llama-swap/llama.cpp config for Qwen3.6 35B-A3B:
qwen3.6-35b-a3b:name:"Qwen3.6 35B-A3B (Coding)"proxy:"http://127.0.0.1/:$%7BPORT%7D"# If you're seeing a `/` after `127.0.0.1` here, don't include it. I think something in Lemmy is trying to "sanitize" this input by adding the `/`.cmd:|
llama-server
--port ${PORT}
--no-webui
-hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL
--jinja
--parallel 1
--flash-attn on
--no-mmproj
--load-mode none
--reasoning-preserve
--ctx-size 190000
--temp 0.6
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.01
A few notes about this config:
Now that I think of it, --reasoning-preservemight be another thing that helped with tool calls.
You can also omit --no-mmproj if you need vision, but it might mean sacrificing speed or context size, so I usually just enable vision in a separate llama-swap model entry to use as needed.
Unsloth recommends --repeat-penalty1.0, but I saw the LLM enter a thinking loop in VS Code, so I bumped it up just a tiny bit to 1.01. I have since seen it do something that resembled the same thought loop, but it was able to recover on its own. Not sure if it’s a coincidence or if 1.01 was actually the solution, so worth some experimentation.
Regarding containerization, familiarize yourself with Dev Containers, they’re super useful for limiting agents to your codebase, with the added bonus that any project you work on comes out of the box with the right version of the tools you need.
They’re oversold though, especially prompt caching and the parameter count war
The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.
Considering which country is better at building power stations, that’s a fascinating dichotomy.
The US going for brute force when the brute force is more available in China… Priceless irony.
China doesn’t have the near the amount of compute resources, but they can run less efficient servers for cheaper, so it’s a wash.
The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.
(It’s not super complex code, just some scripts that I would not have taken to time to write manually.)
I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬
Show me your set up!
There’s not really anything interesting to show. It’s just a home server in a 13 year old desktop ATX case.
There’s no desk, monitor, keyboard, or mouse… But also no cool server rack.
Function over form, and it sits in a spare bedroom out of sight.
EDIT: I found the receipt for the case. It’s a Cougar Volant Black Steel mid tower, purchased in 2013. So my server just looks like this:
Gotta love the black slab.
I meant your LLM stack lol. I just have an RX 6800 XT in my main Linux PC for inference, but it has to share VRAM with the DE. Maybe I’ll set it up for remote development from my laptop instead to free up VRAM.
What are you using? vLLM? llama.cpp? Which params? How much CPU offloading? Do you use draft models? Is it a MoE model? Have you tried llama-swap? Which agentic front-end are you using? I presume you set it up to access it without SSH’ing into the machine, did you do anything special or is it just a raw unsecured open port on the machine to the LAN?
Ohhh lol. Yeah it’s Llama-swap, running llama.cpp for now, but might add vLLM to the llama-swap config to experiment with NVFP4.
I mainly use MoE models so I can get decent speed while using a 150-200k context window. My go-to model has been Qwen3.6 35B-A3B for a while. I tried Qwen3.8 27B, but it was too slow.
Gemma4 26B-A4B also runs nice and fast, but I generally get better results from Qwen3.6. I don’t remember exactly how much CPU offloading is happening, but it’s not much. As long as I can get like 40-50 tokens/sec, I’m usually satisfied enough.
For the coding harness, I’ve been running Pi in an Apple Container (sort of like Podman, but better isolation in a microvm). Though, I recently configured VS Code to use LLMs on my server, and it was actually pretty decent. Still need to explore a bit more, but so far VS Code’s AI capabilities seem much better than they were a year ago (they seemed way behind, back then).
Also, I don’t connect any harness directly to llama-swap. I have another container running Caddy, which acts as a gateway to AI providers. For other services (e.g. OpenRouter), the API key is injected in the Caddy container. I don’t like having API keys or secrets anywhere where LLMs can read them. It’s not so bad for my own self-hosted LLMs, but not cool to send secrets to a server owned by someone else.
How has tool use been for you? I struggled a lot with tool use with Gemma and Qwen, to the point where I needed to build a healing layer.
Regarding the coding harness, I was looking for something CLI-based or JetBrains-based, and I haven’t had much luck getting my local llama.cpp models playing ball with OpenCode. They keep losing context and misusing tools.
I’m not too familiar with Apple containers as I’m running a full Linux stack, but I’ll give Pi a try, seems interesting! Does it work for coding tasks or is it strictly an “orchestrator”?
Tool use with Gemma has been hit or miss. I wouldn’t rely on it for anything unsupervised.
Tool use for Qwen3.6 has been great lately, but I do remember seeing some issues with it too, a while back. I don’t remember when/why the issues cleared up (I have tweaked configs a bit over time), but switching to Pi definitely helped.
I do remember having a lot more problems in OpenCode and it was practically unusable (which is why my recent experience with VS Code was surprising). I’d definitely recommend trying Pi.
A fresh Pi install is very minimal by design. The system prompt is tiny, so it’s a pretty good fit for small LLMs like these. It’s sort of like Neovim: Nothing fancy out of the box, but you can add lots of fancy things to it. I containerize it because I don’t like giving LLMs (especially these small ones) unrestricted access to my host computer – though, I have not seen any signs of it accidentally doing something destructive, which is surprising.
There are similar alternatives to Apple Container for Linux (e.g. Docker Sandboxes, muvm, Firecracker). There’s also this thing made specifically for Pi called Gondolin. I haven’t tried it yet, but I may end up switching to that if it could simplify my stack.
Here’s my current llama-swap/llama.cpp config for Qwen3.6 35B-A3B:
qwen3.6-35b-a3b: name: "Qwen3.6 35B-A3B (Coding)" proxy: "http://127.0.0.1/:$%7BPORT%7D" # If you're seeing a `/` after `127.0.0.1` here, don't include it. I think something in Lemmy is trying to "sanitize" this input by adding the `/`. cmd: | llama-server --port ${PORT} --no-webui -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL --jinja --parallel 1 --flash-attn on --no-mmproj --load-mode none --reasoning-preserve --ctx-size 190000 --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.01A few notes about this config:
--reasoning-preservemight be another thing that helped with tool calls.-MTPpart of the-hfparam. MTP helps speed things up. Here’s the Huggingface page for this model--no-mmprojif you need vision, but it might mean sacrificing speed or context size, so I usually just enable vision in a separate llama-swap model entry to use as needed.--repeat-penalty 1.0, but I saw the LLM enter a thinking loop in VS Code, so I bumped it up just a tiny bit to1.01. I have since seen it do something that resembled the same thought loop, but it was able to recover on its own. Not sure if it’s a coincidence or if1.01was actually the solution, so worth some experimentation.Cheers, I’ll give this a try!
Regarding containerization, familiarize yourself with Dev Containers, they’re super useful for limiting agents to your codebase, with the added bonus that any project you work on comes out of the box with the right version of the tools you need.
eh, just
systemctl isolate multi-user.targetCorrect, that’s how I would do it, but then I need another machine to act as a head.
If your MB has onboard graphics, maybe you could mask the GPU and just pass it off to a container running the LLMs I guess
No onboard graphics unfortunately, that would have been the easy way out.
Reads like American vs European/Asian cars
Always ready to try brute force first. And then some other, less palatable options if that doesn’t work, like slightly less brute force.