DeepSeek V4.1 Flash

Sep
15
DeepSeek V4.1 Flash banner showing a 552B MoE backbone, 8B active on input, 16B on output, a 1M-token context, and $0.003/M off-peak cached input.

DeepSeek V4.1 Flash is a 552B model that bills like a smaller one

DeepSeek V4.1 Flash activates 8B parameters per input token and prices cached input from $0.003 per million. On Terminal-Bench 4.0, it scores 31.2 to Opus 5.0's 51.8.
6 min read