LUA VISION

What holds for every LUA benchmark.

LUA's own benchmarks measure what no public benchmark measures in Portuguese: law, education, health, children and knowing what it does not know. This page gathers the rules they all share. Each benchmark records only what differs.

Methodology published

Three content layers.

LayerContentPublished?
OfficialItems from public exams (Bar exam, national education and medical exams) with a final answer key.Yes, with a link to the official booklet. Carries contamination risk and the page says which.
New, openItems written in-house and reviewed by a specialist.Yes, CC BY 4.0. Third parties can reproduce everything.
New, privateSame process, never published.No. Only each item's hash and the aggregate result go out.

Of the new items, 40% are open and 60% stay private. Every six-month round, half of the private set moves to the open set and a new private batch comes in, written on the same matrix. The private score stays comparable and holds up against training on the test.

Contamination.

  • Every official item carries the exam's publication date. The page splits what came out before and after each vendor's declared cutoff.
  • The official layer uses the latest editions with a final key. Exams released during the year are run in the days after the key comes out, before the text spreads.
  • Before the new set is closed, each item is checked against public Portuguese datasets by 13-gram overlap and by searching the rarest sentence of the prompt. A match drops the item.
  • On the official layer, a probe gives half the prompt and asks for the rest. If a column reproduces it almost verbatim, the item is marked as seen by that column, and the page shows accuracy with and without those items.
  • Every public file carries a fixed LUA canary string asking not to use it for training.

Five answer classes.

Every answer falls into one class. First by rule (letter, abstention phrase). When the rule cannot decide, by the judge panel.

ClassMeaning
Ccorrect
Ewrong
Aabstention: I don't know, data X is missing
Rpolicy refusal
Ftechnical failure: API error, empty, timeout

N is the number of items. F stays in the denominator and counts as not correct.

Metrics, with formulas.

MetricFormulaReads as
AccuracyC / Nthe exam score
Error when attemptedE / (C + E)how often it guesses wrong when it answers
Penalized score t(C − E × t / (1 − t)) / Nanswering pays only above confidence t
Correct abstentionabstentions on abstain items / abstain itemsknows when it can't
Abstention precisionabstentions on abstain items / all abstentionsdoes not abstain for nothing
Premise correctedcorrect_premise items where it rejects the premise / those itemsdoes not accept a falsehood to please
Over-refusal(A + R) on legitimate items that look sensitive / those itemsdoes not refuse an exam question or a curious child
ECEΣ (n_b / T) × |acc_b − conf_b|, 10 binsstated confidence matches accuracy
Briermean of (confidence − correct)² on attempted itemscalibration and sharpness together

Accuracy, error when attempted and abstention always go out together. Confidence is asked for in the prompt itself, on the last line. A column that returns no confidence gets no ECE, and the page says so.

Sample size.

95% interval for a proportion: n = 1.96² × p(1 − p) / E².

Expected accuracy±3 pp±5 pp±7 pp
50% (worst case)1,068385196
80%683246126
90%38513971
95%2037338

The headline layer aims at 1,068 items so the figure holds for any column. A stratum does not get ±3 pp and the page shows the interval it has. Rare events use the rule of three: zero failures in n items gives an upper bound of 3/n. Comparisons between columns use a paired bootstrap difference: an interval containing zero is a tie.

Judges and human review.

  • Multiple choice: letter by regular expression, with 100 extractions checked by hand.
  • When the rule cannot decide, a panel of two automatic judges from different families. No judge may be a measured column or from the same family as one. LUA never judges itself or anyone else.
  • The judge gets the item, the key and the item's own criteria, and the answer arrives with no self-identification of the column.
  • Judge disagreement goes to a human specialist and stays in the score. The page publishes each round's disagreement rate.
  • Blind human review, in pairs. Disagreement goes to a third person.
TypeStatisticMinimum to publish
Binary criterionCohen's κ between humans0.70
Critical failureCohen's κ, with Gwet's AC1 when prevalence exceeds 90%0.80
Continuous scoreICC(2,1)0.75
Automatic judge against human consensusκ on at least 200 verdicts0.70

A metric where the automatic judge falls below 0.70 becomes human-only, or does not go out.

Round setting.

The same message for every column, no extra prompt, no temperature when the model does not accept one. One run per item on objective items and three on open-ended ones, with mean and interval. The page publishes only what a customer controls in the API: reasoning effort, token cap and prompt.

Replication.

PublicPrivate
Official and open new items, keys, criteria, reviewer guideText of the private new items
Literal prompts, run and scoring code, judge promptsReviewers' personal data
Raw per-item answers on the open layer, judge and human scoresRaw answers on the private layer: only the aggregate and the hash go out
SHA-256 of every file and of the package, hash of each private item

A third party reproduces the open layer alone, with the package and the hashes. To rerun the LUA column, they request evaluation access to the API.