More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Laguna XS 2.1 packs 33 billion parameters in a Mixture-of-Experts setup, activating 3 billion per token. It keeps the XS.2 architecture but jumps from 57.7% to 63.1% on SWE-bench Multilingual, outpacing Qwen 3.6 (35B) and North Mini Code (30B). It also posts stronger scores on Verified, Pro and Terminal-Bench 2.0, making it a sharper tool for agentic coding and long-horizon tasks on local machines.
You can run XS 2.1 under vLLM, SGLang, NVIDIA TensorRT-LLM, Hugging Face Transformers and Ollama, with llama.cpp support coming soon. Three quantized checkpoints (FP8, INT4 and NVFP4) let you squeeze into lower-VRAM rigs. Add the DFlash speculator models we trained for each checkpoint and you’ll double token throughput versus stock XS 2.1. We serve the model at 256 K context length via our API and OpenRouter.
The code’s under OpenMDW-1.1, a fully permissive license co-sponsored by NVIDIA and the Linux Foundation. You can download BF16, FP8, NVFP4 and INT4 weights from Hugging Face, spike the model on OpenRouter (poolside/laguna-xs-2.1) or hit our API at the same rates as XS.2: $0.10 input, $0.20 output and $0.05 cache-read per million tokens. Ollama, llama.cpp, TRT-LLM and vLLM all work locally—just drop in the DFlash draft for a speed boost.
Laguna XS.2 will leave our API in one week but stays on Baseten’s Model Library. We want your feedback: compare XS 2.1 side by side with XS.2 and let us know where it shines or stumbles. Hit up our Discord, email models@poolside.ai or ping us on X.
Questions about this article
No questions yet.