# Pedro Cuenca runs MiMo 2.6 Flash across RTX 6000 and M5 laptop via llama.cpp

Pedro Cuenca demonstrated running MiMo 2.6 Flash with native mxfp4 weights across an RTX 6000 GPU and an M5 laptop over 10 GbE at 40 tokens per second, utilizing out-of-the-box support in llama.cpp.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/bbb925c04bfb4271b3d0b3ae.html

Content language: en
Localization state: sameLanguage

Source: Pedro Cuenca · Published 10/7/2026, 11:03:30

Pedro Cuenca demonstrated running MiMo 2.6 Flash across heterogeneous hardware by combining an RTX 6000 GPU and an M5 laptop connected over 10 GbE. The setup achieved an inference speed of 40 tokens per second.

The execution ran the model's native mxfp4 weights directly and is supported out of the box in llama.cpp.

Tags: llama.cpp, MiMo 2.6 Flash, Inference

[View original post](https://x.com/pcuenq/status/2107825758442082591)
