The metric that picks your next model appears to be cost per task, not benchmark rank — type0 | type0