Most clinical-AI benchmarks answer the wrong question. A new Python library tries to fix that. — type0 | type0