people · May 16, 2026
E2B Sandboxes Integrated into Gemma 4 E2B Local AI Model Testing
Share the canonical public link.
Witcheer tested Gemma 4 E2B model on RTX 4060 Ti 8GB hardware on May 14, 2026, achieving 117.8 tokens per second decode speed. The 2.3 billion active parameter model fits in 2.6 GB VRAM and runs at 2.9 GB on disk. Gemma 4 E2B outperformed larger models like Mistral 7B in speed while maintaining context stability up to 65K tokens. The testing compared it directly to Qwen3.6 35B and GLM 4.7 Flash on the same rig.
Spend governor blocked model creation: provider_circuit_open (lane=dev, provider=together)