logo

KnowBench: Evaluating clinical AI with effort reduction

Posted by kangjl888 |2 hours ago |1 comments

kangjl888 2 hours ago

Knowtex (YC S22) is the first frontier AI Lab for healthcare. We build AI that automates clinical work and provides clinical intelligence to providers. Standard evals in this space use medical exams or use expert rubric panels. These do not measure whether a clinician's workload actually was decreased.

Our observation is that this measure already exists in the workflow. Before a note enters the medical record, a clinician reviews the draft and fixes whatever is wrong, then signs and owns it legally. The edits represents the measurement in that whatever survives review is effort Knowtex reduced, and whatever got edited by the clinician is effort returned.

We formalized this as our benchmark KnowBench measuring Effort Reduction: the fraction of generated content accepted without content edits. The same construction works for billing codes, orders, and chart summaries, not just notes. Our arXiv paper defines the metric, its failure modes (you can game it with short drafts, clinicians can rubber-stamp, etc.), and a six-item reporting checklist.

Our result showed 97.99% aggregate across 1M+ signed encounters and 13 specialties.

Closest prior art is HTER from machine translation and Copilot's acceptance rate, with one difference: our reviewer signs the output into a legal record.

Paper: https://arxiv.org/abs/2609.15794