BL-SAU: exam, protocol and warning signs.
S1 Exam: medical knowledge at the level of Brazilian exams.
Methodology publishedMethodology published · round in preparation
NoticeThis benchmark is not clinical validation. It does not certify clinical use and does not replace a medical assessment.
Methodology published · round in preparation.
No result figures on this page. The benchmark has not been run yet. The sizes below are the sample design, not results.
Capability measured.
S2 Care: decides according to the Ministry of Health protocol in force.
S3 Layperson: recognizes warning signs and refers, without diagnosing.
S4 Over-refusal: answers a legitimate clinical question that looks sensitive. It exists because of a real finding: on the Brazil Bench, two legitimate ENAMED questions were refused or blocked in more than one column. Refusing a medical exam question is an error, and the benchmark shows it.
Use cases.
- Health plans and hospitals putting an assistant on message triage.
- Clinics using an assistant for post-procedure guidance.
- Medical teams looking up protocols.
- The decision at stake is whether patient guidance can go out without a doctor in the loop, and in which cases never.
Sources.
| Layer | Content | Use |
|---|---|---|
| Official | Revalida objective stage, INEP exams and final keys. ENAMED is already on the Brazil Bench and is not repeated | S1 |
| Regulatory | Ministry of Health clinical protocols, therapeutic guidelines and management guides, published by ordinance | S1 and S2 |
| New | Questions and cases written from the protocols, citing the passage and ordinance version in the item | S1 to S4 |
No real patient data. Every case is fictional, with no name, photo or real test result.
Format and planned size.
| Row | Content | n planned |
|---|---|---|
| S1 Exam | about 400 official and 700 new (280 open, 420 private), across the five main areas | 1,100 |
| S2 Care | open-ended cases with required and forbidden criteria, such as allowing anti-inflammatories when dengue is suspected | 150 |
| S3 Layperson | lay messages: refer urgently, refer non-urgently or answer; half with a warning sign | 300 |
| S4 Over-refusal | legitimate clinical questions asked as a professional or student | 150 |
Metrics.
- S1: accuracy, error when attempted, penalized score, ECE
- S2: share of required criteria met; a critical failure is any forbidden criterion present, published as a count with the 95% upper bound, never folded into an average
- S3: sensitivity to warning signs, specificity, and critical failure when a warning sign is present and the answer does not refer
- S4: over-refusal, side by side with S3, because refusing everything zeroes critical failures and inflates over-refusal
Judging and human review.
- S1 by rule.
- S2 and S3: a panel of two automatic judges; two doctors with active registration blindly review 30% of each column and 100% of critical failures flagged by any judge.
- S4: a doctor decides every refusal.
- Every new item has one doctor as author and another as reviewer. κ target on critical failure: 0.80.
Replication.
- Public: official and open S1, open S2 and S3 with criteria, full S4.
- Private: new parts of S1, S2 and S3.
Risks and limits.
- Protocols change. An updated protocol voids items, and each item has an expiry date.
- The line between correct and wrongful refusal in health (dose, overdose, self-harm) is a house policy decision and must be written before the first S4 item.
- This benchmark's text goes through regulatory review before release, because clinical performance in commercial material can be read as a claim about health software.