Did Codex Reset
GitHub

Pedro Cuenca runs MiMo 2.6 Flash across RTX 6000 and M5 laptop via llama.cpp

Pedro Cuenca

Pedro Cuenca demonstrated running MiMo 2.6 Flash across heterogeneous hardware by combining an RTX 6000 GPU and an M5 laptop connected over 10 GbE. The setup achieved an inference speed of 40 tokens per second.

The execution ran the model's native mxfp4 weights directly and is supported out of the box in llama.cpp.