Thanks for testing it out and sharing these numbers! 🙌
"Run a swarm of these locally" — that's exactly the vision we had. 2B params, 1.56GB Q4, 200+ tok/s on a 4090, and still keeping up with 4B models. That's what local AI agents should look like.
Can't wait to see what you build with them 🔥 https://x.com/outsource_/status/2097005719983689998/video/1