LUA VISION

BL-SAU: exam, protocol and warning signs.

S1 Exam: medical knowledge at the level of Brazilian exams.

Methodology publishedMethodology published · round in preparation
NoticeThis benchmark is not clinical validation. It does not certify clinical use and does not replace a medical assessment.

Methodology published · round in preparation.

No result figures on this page. The benchmark has not been run yet. The sizes below are the sample design, not results.

Capability measured.

S2 Care: decides according to the Ministry of Health protocol in force.

S3 Layperson: recognizes warning signs and refers, without diagnosing.

S4 Over-refusal: answers a legitimate clinical question that looks sensitive. It exists because of a real finding: on the Brazil Bench, two legitimate ENAMED questions were refused or blocked in more than one column. Refusing a medical exam question is an error, and the benchmark shows it.

Use cases.

  • Health plans and hospitals putting an assistant on message triage.
  • Clinics using an assistant for post-procedure guidance.
  • Medical teams looking up protocols.
  • The decision at stake is whether patient guidance can go out without a doctor in the loop, and in which cases never.

Sources.

LayerContentUse
OfficialRevalida objective stage, INEP exams and final keys. ENAMED is already on the Brazil Bench and is not repeatedS1
RegulatoryMinistry of Health clinical protocols, therapeutic guidelines and management guides, published by ordinanceS1 and S2
NewQuestions and cases written from the protocols, citing the passage and ordinance version in the itemS1 to S4

No real patient data. Every case is fictional, with no name, photo or real test result.

Format and planned size.

RowContentn planned
S1 Examabout 400 official and 700 new (280 open, 420 private), across the five main areas1,100
S2 Careopen-ended cases with required and forbidden criteria, such as allowing anti-inflammatories when dengue is suspected150
S3 Laypersonlay messages: refer urgently, refer non-urgently or answer; half with a warning sign300
S4 Over-refusallegitimate clinical questions asked as a professional or student150

Metrics.

  • S1: accuracy, error when attempted, penalized score, ECE
  • S2: share of required criteria met; a critical failure is any forbidden criterion present, published as a count with the 95% upper bound, never folded into an average
  • S3: sensitivity to warning signs, specificity, and critical failure when a warning sign is present and the answer does not refer
  • S4: over-refusal, side by side with S3, because refusing everything zeroes critical failures and inflates over-refusal

Judging and human review.

  • S1 by rule.
  • S2 and S3: a panel of two automatic judges; two doctors with active registration blindly review 30% of each column and 100% of critical failures flagged by any judge.
  • S4: a doctor decides every refusal.
  • Every new item has one doctor as author and another as reviewer. κ target on critical failure: 0.80.

Replication.

  • Public: official and open S1, open S2 and S3 with criteria, full S4.
  • Private: new parts of S1, S2 and S3.

Risks and limits.

  • Protocols change. An updated protocol voids items, and each item has an expiry date.
  • The line between correct and wrongful refusal in health (dose, overdose, self-harm) is a house policy decision and must be written before the first S4 item.
  • This benchmark's text goes through regulatory review before release, because clinical performance in commercial material can be read as a claim about health software.