DeepSeek V4 doesn't bolt vision on. It rewrites each image as 384 tokens that share the model's main text sequence. — type0 | type0