- Sat 03 October 2026
- Linux
Self-Hosted AI: Qwen 3.8 on a Rented Blackwell at 150 Tokens per Second
An open-weight Qwen 3.8 27B model, an RTX PRO 6000 Blackwell rented by the hour, and one shell command to start it. For most of my everyday coding it is good enough, it answers at roughly 150 tokens per second, and my prompts stay on infrastructure I chose. Some honest caveats about what “self-hosted” means on someone else’s GPU are included, along with the workflow that has worked best for me: a frontier model for architecture, and a fast open-weight model for the token-heavy implementation loop.