LUA VISION

BL-EDU: Brazilian school content.

LUA's own benchmark on Brazilian school content: objective questions from the 2024 and 2025 national high-school exam (ENEM), with the official INEP answer key. Only the LUA Genesys PI House model was measured, no competitor in this round.

RunLUA Record · run by LUA Vision · not an independent evaluation

Result.

CutnAccuracy (%)95% CIError when attempted
Overall, 2024+202535499.798.4–100.00.3

By area.

AreanAccuracy (%)95% CI
Linguagens9098.994.0–99.8
Ciências Humanas90100.095.9–100.0
Ciências da Natureza86100.095.7–100.0
Matemática88100.095.8–100.0

By edition.

EditionnAccuracy (%)95% CI
2024179100.097.9–100.0
202517599.496.8–99.9

Image coverage: 125 of 125 questions with an image, chart or table were run, all through the INEP's official audio-description transcript, with 100.0% accuracy.

The round's only mistake: a 2025 Language Arts question where the answer came with a stated confidence of 97, against the official key.

95% CI by Wilson. Score with penalty t=0.5: 99.6.

Methodology.

ModelLUA Genesys PI House, public API api.lua.vision, medium effort.
DataOfficial ENEM 2024 and 2025 booklets (version with image transcribed to text) and the INEP's final printed answer key.
ScoringLetter extracted by regular expression, no automatic judge.
RunsOne per item, pass@1.

Contamination.

No declared training cutoff exists for the evaluated model, so this page does not estimate one. Both editions used (2024 and 2025) have been public for over ten months, plenty of time to enter any web-collected corpus — both are contamination candidates, not just the older one. Accuracy by edition did not drop on the more recent exam (2024: 100%; 2025: 99.4%, overlapping intervals), which does not point to memorization concentrated in one of the two, but it also does not separate "a very good model" from "both were seen in training". No memorization probe was run in this delivery. The 99.7% accuracy is high even for a widely discussed public exam, and should be read as this round's result, not as a capability confirmedly free of contamination.

What did not run.

  • Essay: not executed. The official five-competency matrix sits in a PDF with a font that has no readable character table; two different extractors returned inconsistent text, and using a matrix rebuilt from memory would mean inventing the grading criteria. 100% human review by two teachers was also unavailable in this round.
  • Tutor: not executed. It would measure whether LUA points out a student's mistake without handing over the finished answer, but it would need a real student-error case; inventing a fictional one would fabricate the source data, which the benchmark exists to avoid.

Limitations.

  • Measures whether LUA gets exam questions right, not whether it teaches. A high score here does not license saying the product teaches well.
  • N per area sits between 86 and 90, so the per-area confidence interval is wide, 4 to 6 points; only the two-edition aggregate has the narrower interval.
  • The IRT-estimated proficiency uses INEP's published parameters outside a real sitting; it is not comparable to a real candidate's ENEM score, and at the Math extreme (100% accuracy) the estimate extrapolates beyond the region where the calibration has the most reference candidates.
  • Language coverage: only the English option of the foreign-language questions.

Package and files.

bl-edu-2026-09.tar.gz245ee3af59e2b1c4a06ab37ce27dd0ace58045d5351cb4473a7bfc56fcc9d0cc
FileSizeDownload
bl-edu-2026-09.tar.gz31 KBDownload
Per-item resultSHA-256Download
resultados/e1_objetivas.house.jsonl2d842b0ad37aa5eeaba1ab35c9a682f6a8f64789c2bec7b15fc651a9a851c57aDownload