Artificial Analysis open-sources AA-AgentPerf-Local for laptop and workstation inference
Artificial Analysis is open-sourcing the code and data for AA-AgentPerf-Local, an inference test for local AI models on laptops and workstations, alongside a leaderboard of model and hardware combinations. Its default workload replays 8 recorded agentic tasks across 168 model turns. Each request includes the full conversation so far, growing context to about 56K tokens, and every system generates the recorded token count. Tool execution is skipped by default to isolate inference speed; users can instead replay recorded tool delays or run tool calls live. The launch measures one agent using the whole system, with multi-agent and shared-workload coverage planned later.
The initial hardware is the NVIDIA DGX Spark (128 GB), AMD Ryzen AI Halo (128 GB), MacBook Pro M5 Pro (64 GB), and NVIDIA GeForce RTX 5090 (32 GB), covering CUDA, ROCm, Vulkan, and Metal. Featured models, all at 4-bit, are Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B total, 5B active). The tool can also test any model served by an OpenAI-compatible inference server. All 14 published serving configurations use speculative decoding through MTP, DFlash, or DSpark, using an official configuration where one existed.
Artificial Analysis says Qwen3.6-35B-A3B, with 3B active parameters, was the fastest model on every system, 2.5–3.3 times faster than dense Qwen3.8-27B, while Ling 3.0 Flash still finished behind Qwen3.5-9B on hardware that could run it. For every model that fit in 32 GB, the RTX 5090 was fastest, with completion times more than 3.5 times shorter than the other systems; its 1,792 GB/s memory bandwidth compares with 256–307 GB/s on the unified-memory systems. DGX Spark and Ryzen AI Halo share 128 GB of unified memory and a $4,000 launch price. Spark was 1.4–1.7 times faster on three of four models and tied on Qwen3.5-9B, a gap Artificial Analysis says exceeds their 7% bandwidth difference. The M5 Pro MacBook, the only laptop tested, has 307 GB/s bandwidth and a current price of $3,700. It finished within 2–9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B, but 21% slower on Qwen3.5-9B.