product · May 15, 2026
Poolside AI Releases Laguna XS.2 Coding Model Optimized for Single GPU Deployment
Share the canonical public link.
Poolside AI team members Pupposandro, Davide Ciffa, and Luceboxai deployed Laguna XS.2 on a single RTX 3090 GPU achieving 111 tokens per second decode speed on May 14, 2026. The model demonstrates 5.4 times faster 128K prefill performance compared to llama.cpp benchmarks. Community developers contributed to making Laguna XS.2 the first mixture-of-experts target for PFlash attention mechanism. Hugging Face Transformers integrated Laguna XS.2 support in version 5.7.0 released May 14, 2026.
Failed after 3 attempts. Last error: Service unavailable