# DeepLearning.AI details DeepSeek cache reduction and Flash benchmark results in The Batch

In an issue of The Batch, DeepLearning.AI analyzed DeepSeek's architecture, reporting an 890-byte cache per token that is 437 times smaller than DeepSeek-V1. The analysis also notes a 25% compute increase per output token from 4K to 1M input tokens and highlights Flash outperforming V4-Pro on the AA Index.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/f37d751491a64ac5a3847135.html

Content language: en
Localization state: sameLanguage

Source: DeepLearning.AI · Published 10/7/2026, 17:02:37

DeepLearning.AI published a [technical breakdown](https://hubs.la/Q04ztt000) of DeepSeek in this week's issue of The Batch, focusing on memory footprint optimizations for AI agents. The report states that DeepSeek reduced its cache to 890 bytes per token, making it 437 times smaller than DeepSeek-V1.

The analysis also notes a 25% increase in compute per output token when scaling context from 4K to 1M input tokens. On benchmark evaluation, Flash scored 39 compared to 36 for V4-Pro on the AA Index, with costs reported at $0.27 per task versus $0.67 per task.

Tags: DeepLearning.AI, DeepSeek, Benchmarks, LLM Inference

[View original post](https://x.com/DeepLearningAI/status/2107984315598434418)
