What holds for every LUA benchmark.
LUA's own benchmarks measure what no public benchmark measures in Portuguese: law, education, health, children and knowing what it does not know. This page gathers the rules they all share. Each benchmark records only what differs.
Three content layers.
| Layer | Content | Published? |
|---|---|---|
| Official | Items from public exams (Bar exam, national education and medical exams) with a final answer key. | Yes, with a link to the official booklet. Carries contamination risk and the page says which. |
| New, open | Items written in-house and reviewed by a specialist. | Yes, CC BY 4.0. Third parties can reproduce everything. |
| New, private | Same process, never published. | No. Only each item's hash and the aggregate result go out. |
Of the new items, 40% are open and 60% stay private. Every six-month round, half of the private set moves to the open set and a new private batch comes in, written on the same matrix. The private score stays comparable and holds up against training on the test.
Contamination.
- Every official item carries the exam's publication date. The page splits what came out before and after each vendor's declared cutoff.
- The official layer uses the latest editions with a final key. Exams released during the year are run in the days after the key comes out, before the text spreads.
- Before the new set is closed, each item is checked against public Portuguese datasets by 13-gram overlap and by searching the rarest sentence of the prompt. A match drops the item.
- On the official layer, a probe gives half the prompt and asks for the rest. If a column reproduces it almost verbatim, the item is marked as seen by that column, and the page shows accuracy with and without those items.
- Every public file carries a fixed LUA canary string asking not to use it for training.
Five answer classes.
Every answer falls into one class. First by rule (letter, abstention phrase). When the rule cannot decide, by the judge panel.
| Class | Meaning |
|---|---|
| C | correct |
| E | wrong |
| A | abstention: I don't know, data X is missing |
| R | policy refusal |
| F | technical failure: API error, empty, timeout |
N is the number of items. F stays in the denominator and counts as not correct.
Metrics, with formulas.
| Metric | Formula | Reads as |
|---|---|---|
| Accuracy | C / N | the exam score |
| Error when attempted | E / (C + E) | how often it guesses wrong when it answers |
| Penalized score t | (C − E × t / (1 − t)) / N | answering pays only above confidence t |
| Correct abstention | abstentions on abstain items / abstain items | knows when it can't |
| Abstention precision | abstentions on abstain items / all abstentions | does not abstain for nothing |
| Premise corrected | correct_premise items where it rejects the premise / those items | does not accept a falsehood to please |
| Over-refusal | (A + R) on legitimate items that look sensitive / those items | does not refuse an exam question or a curious child |
| ECE | Σ (n_b / T) × |acc_b − conf_b|, 10 bins | stated confidence matches accuracy |
| Brier | mean of (confidence − correct)² on attempted items | calibration and sharpness together |
Accuracy, error when attempted and abstention always go out together. Confidence is asked for in the prompt itself, on the last line. A column that returns no confidence gets no ECE, and the page says so.
Sample size.
95% interval for a proportion: n = 1.96² × p(1 − p) / E².
| Expected accuracy | ±3 pp | ±5 pp | ±7 pp |
|---|---|---|---|
| 50% (worst case) | 1,068 | 385 | 196 |
| 80% | 683 | 246 | 126 |
| 90% | 385 | 139 | 71 |
| 95% | 203 | 73 | 38 |
The headline layer aims at 1,068 items so the figure holds for any column. A stratum does not get ±3 pp and the page shows the interval it has. Rare events use the rule of three: zero failures in n items gives an upper bound of 3/n. Comparisons between columns use a paired bootstrap difference: an interval containing zero is a tie.
Judges and human review.
- Multiple choice: letter by regular expression, with 100 extractions checked by hand.
- When the rule cannot decide, a panel of two automatic judges from different families. No judge may be a measured column or from the same family as one. LUA never judges itself or anyone else.
- The judge gets the item, the key and the item's own criteria, and the answer arrives with no self-identification of the column.
- Judge disagreement goes to a human specialist and stays in the score. The page publishes each round's disagreement rate.
- Blind human review, in pairs. Disagreement goes to a third person.
| Type | Statistic | Minimum to publish |
|---|---|---|
| Binary criterion | Cohen's κ between humans | 0.70 |
| Critical failure | Cohen's κ, with Gwet's AC1 when prevalence exceeds 90% | 0.80 |
| Continuous score | ICC(2,1) | 0.75 |
| Automatic judge against human consensus | κ on at least 200 verdicts | 0.70 |
A metric where the automatic judge falls below 0.70 becomes human-only, or does not go out.
Round setting.
The same message for every column, no extra prompt, no temperature when the model does not accept one. One run per item on objective items and three on open-ended ones, with mean and interval. The page publishes only what a customer controls in the API: reasoning effort, token cap and prompt.
Replication.
| Public | Private |
|---|---|
| Official and open new items, keys, criteria, reviewer guide | Text of the private new items |
| Literal prompts, run and scoring code, judge prompts | Reviewers' personal data |
| Raw per-item answers on the open layer, judge and human scores | Raw answers on the private layer: only the aggregate and the hash go out |
| SHA-256 of every file and of the package, hash of each private item |
A third party reproduces the open layer alone, with the package and the hashes. To rerun the LUA column, they request evaluation access to the API.