Background-work AI runs in the back office of every enterprise, and the unit of accounting is changing. Invoice processing, document review, and vulnerability remediation are loops that burn through tokens by the thousand while a human never sees the output. The metric that decides which model ships to production appears to be no longer a benchmark score. It is becoming the cost of completing one task end to end.
Google's blog framed Flash-Lite as its fastest 3.5-class model, hitting 350 output tokens per second on the Artificial Analysis Index. That number is a procurement signal, not a vanity speed: throughput per dollar on high-volume, low-stakes background work is the dimension on which this tier competes. The 17 percent output-token drop Google reports for 3.6 Flash versus 3.5 Flash, and the up-to-65 percent drop it cites on the Datacurve DeepSWE coding test, point the same way. The contest is cost-per-task, not leaderboard rank. One product line became three SKUs because the cost curve no longer has a single shape.
The vendor-aligned benchmarks and the named showcase customers (Figma, Hebbia, Harvey) deserve the usual discount, and the restricted-access Flash Cyber variant has not been independently assessed. The mechanism is what travels: read every future model release as a price-per-task number, ask which background workload it prices, and decide procurement on tokens burned per resolved ticket, not on the test suite it topped.
Reported by Sky for Type0, from Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Read the original: blog.google