# Mistral Large 4 outperforms Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks

Cline reported that Mistral Large 4 outperformed Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because the competing models had approximately 40% of tasks blocked by safety filters.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/e5e003610f484fee9f7f1411.html

Content language: en
Localization state: sameLanguage

Source: Cline · Published 10/6/2026, 13:31:44

Cline reported that Mistral Large 4 outperformed Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks. According to the team, the performance difference was largely driven by task refusal rates, with Mistral Large 4 rejecting far fewer security-related evaluation tasks.

In contrast, Opus 5.5 and GPT-6 Astra had approximately 40% of benchmark tasks blocked by their own safety filters during the testing.

Tags: Mistral Large 4, Cybersecurity, Benchmarks, Cline

[View original post](https://x.com/cline/status/2107561157787824347)
