LUA VISION

BLUEX: Unicamp and USP entrance exams.

Questions from the Unicamp and USP university entrance exams, across every high-school subject. It measures Brazilian school knowledge and reading of long prompts in Portuguese.

LUA Record · run by LUA Vision · not an independent evaluation

Results.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

ModelAccuracy (%)
↑ higher is better
95% CIAnsweredRefused or blockedNo answer
LUA Genesys PI House96.094.6–97.380270
GPT-5.596.394.9–97.580270
GPT-5.495.493.9–96.880270
Grok 4.696.394.9–97.580513
DeepSeek-V4-Pro93.892.1–95.480720
Kimi K2.694.192.5–95.6797120
Sabiá 4 Thinking†93.0————

Accuracy (%) · higher is better

GPT-5.5
96.3IC 94.9–97.5
Grok 4.6
96.3IC 94.9–97.5
LUA Genesys PI House
96.0IC 94.6–97.3
GPT-5.4
95.4IC 93.9–96.8
Kimi K2.6
94.1IC 92.5–95.6
DeepSeek-V4-Pro
93.8IC 92.1–95.4
Sabiá 4 Thinking †
93.0no CI published

Axis from 90 to 100. Black bar: 95% confidence interval.

How it was run.

Date24/09/2026
Datasetportuguese-benchmark-datasets/BLUEX, commit cb5d60a88a4f, file data/questions-00000-of-00001.parquet
Items809 of 1,422. 613 are out: questions with an associated image or with no A-to-E key, for every model.
ScoringExact letter, extracted by regular expression from the final line "Resposta: X". No judge.
RunsOne per item, pass@1.
FailuresAn unanswered question scores zero and stays in the denominator. That includes refusals by the provider's content filter and refusals by the model itself.
Interval95% interval by percentile bootstrap, 4,000 resamples over items, fixed seed.
SHA-256 of the datadados/bluex.jsonl
d57ecbb34132de8cfd9d99f3e14cb4dae353c8e7ecc13a20fd38324a4cf33362

Setting per model.

Every model got the same messages, through the commercial API of each vendor or of the cloud where the model is published, on the same date. No extra system prompt, no temperature, no tools beyond those the benchmark defines.

ModelAccessReasoning effort
LUA Genesys PI Housepublic API api.lua.visionmedium
GPT-5.5API comercialmedium
GPT-5.4API comercialmedium
Grok 4.6API comercialmodel default
DeepSeek-V4-ProAPI comercialmodel default
Kimi K2.6API comercialmodel default
Sabiá 4 Thinking †figure published by the vendor—

Protocol decisions.

  • A single zero-shot multiple-choice prompt, the same for every model: the model reasons and ends with "Resposta: X". The prompt is in the package.
  • Questions with images are dropped for every model, because not every compared model reads images.

Defects found in the data.

  • Seven questions with literary and documentary text (among them excerpts from the Letter of Pero Vaz de Caminha, from Jorge de Lima and from a UN report on gender bias) were refused in the LUA column and blocked by content filters in other columns. They score zero. The count for each column is in the table.

Contamination.

The items are in public datasets. Any column, LUA included, may have seen these questions in training. This round did not measure contamination, so read the figure as a ceiling on what the model knows about the subject, not as proof of reasoning on unseen items.

Limitations.

  • One run per item. Run-to-run variation of the same model was not measured.
  • A single zero-shot multiple-choice prompt. Another prompt may change the order of the columns.
  • The figure holds for the model served by the API on 24/09/2026. A new version needs a new round.
  • Questions with images were left out for every column. The result says nothing about reading figures.

Reproduce.

The package holds the code that calls the models, builds each test, scores it and builds the table, plus the aggregated table for this round. The data is downloaded by preparar_dados.py at the exact versions and checked by SHA-256. If one byte changes, the script stops.

With Python 3 and Node 20 or newer:

tar -xzf bancada-brasil-2026-09.tar.gz && cd bancada-brasil-2026-09
shasum -a 256 -c <(sed -n '/^Arquivos deste pacote/,/^$/p' VERIFICAR.txt | grep -E '^[0-9a-f]{64}')
pip install huggingface_hub pyarrow
python3 preparar_dados.py          # baixa, normaliza e confere SHA-256
export LUA_API_KEY=...             # chave de avaliação da API LUA
node run.mjs --bench bluex --modelo house
node tabela.mjs
bancada-brasil-2026-09.tar.gz1144df20af11b5c0141614c256fa1733c698bf829be3303842839ec149a1f81111 KB · Raw per-item answers are not in this package.
FileSizeDownload
bancada-brasil-2026-09.tar.gz11 KBDownload
resultados-agregados.json7 KBDownload
VERIFICAR.txt1 KBDownload

History.

24/09/2026 · first round, published on this page.