Mistral AI 稱 Mistral Large 4 在 Harvey 法律智能體基準測試中位居開源模型首位 MMistral AI14小時41分鐘前 Mistral AI 稱,Mistral Large 4 在 Harvey 的法律智能體基準測試(Legal Agent Benchmark,LAB)中名列開源模型第一。 #mistral-ai#mistral-large-4#基準測試
Mistral AI 公布 Mistral Large 4 於 AutomationBench 上的 657 項工作流程測試結果Mistral AI 公布了 AutomationBench 評測結果。在模擬辦公應用的 657 項業務工作流程中,Mistral Large 4 完成了測試,表現領先於 Kimi K3、MiMo-V2.6-Pro 和 DeepSeek V4 Pro。
Mistral Large 4 在網絡安全 CTF 競速賽中攻破 19 項挑戰中的 18 項Mistral AI 表示,Mistral Large 4 在基於實際運行環境的奪旗賽(CTF)競速測試中透過工具調用攻破了 19 項挑戰中的 18 項。
Mistral AI 推出擁有 1 萬億參數的 Mistral Large 4 並即時開放 APIMistral AI 發布原生多模態模型 Mistral Large 4,總參數達 1 萬億,激活參數為 490 億。該模型現已透過 API 及 Mistral Cloud 基礎設施提供,並計劃於 10 月底開放模型權重。
Mistral AI 示範 Mistral Large 4 於 12 分鐘內逆向工程惡意軟件Mistral AI 示範了 Mistral Large 4 端對端分析未知二進制檔案。該模型在 12 分鐘內將樣本識別為 Cobalt Strike,提取了威脅指標(IoC)與惡意軟件配置,並生成包含 YARA 規則的報告。
Mistral Large 4 位列 Code Arena: WebDev 榜單第 45 位Mistral Large 4 以 1,534 分登陸 Arena.ai 的 Code Arena: WebDev 榜單第 45 位,與 Claude Opus 4.8 (High) 差距僅 2 分,混合 Token 成本約為後者的六分之一。