LUA would pass the Bar exam, on both stages, across the 7 most recent editions.
LUA gets 99.7% of the 2024 and 2025 ENEM objective questions right.
354 of 354 questions, 100% image coverage.
See the full exam →LUA Record · run by LUA Vision · not an independent evaluation. No figure was rounded.
Genesys PI evaluations.
LUA Vision is a Brazilian language model lab. We evaluate Genesys PI on public benchmarks and on in-house benchmarks, side by side with frontier models, on the same date and under the same protocol. Every result carries a confidence interval and a SHA-256 sealed package so anyone can redo the math.
Results by capability.
LUA Genesys PI House first, then the comparison models. Each row leads to the benchmark page, with the confidence interval, each model's setting and the files.
In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).
Genesys PI has the best result in the comparison on 7 of the 10 metrics. A tie on the best value counts for every tied model.
| Benchmark | LUA Genesys PI | GPT-5.5 | GPT-5.4 | Grok 4.6 | DeepSeek-V4-Pro | Kimi K2.6 | Sabiá 4 |
|---|---|---|---|---|---|---|---|
| Specialized knowledge in Portuguese | |||||||
| ENAMEDAccuracy (%)↑ higher is better | 96.6best | 96.6 | 94.3 | 95.4 | 94.3 | 93.1 | 94.4† |
| POSCOMPAccuracy (%)↑ higher is better | 95.4 | 95.4 | 95.4 | 96.2 | 93.1 | 95.4 | 90.8† |
| BLUEXAccuracy (%)↑ higher is better | 96.0 | 96.3 | 95.4 | 96.3 | 93.8 | 94.1 | 93.0† |
| Agents and tool use | |||||||
| Ticket-Bench PTAccuracy (%)↑ higher is better | 87.4best | 86.1 | 86.8 | —‡ | 85.4 | 84.8 | — |
| Truthfulness and calibration | |||||||
| SimpleQA VerifiedF1↑ higher is better | 75.2best | 65.5 | 45.9 | 55.3 | 46.2 | 39.9 | — |
| AA-Omniscience-PublicOmniscience Index↑ higher is better | 56.8best | 35.3 | 20.7 | 37.5 | -2.7 | 19.2 | — |
| HalluLens · MixedEntitiesFalse acceptance (%)↓ lower is better | 33.0best | 89.2 | 67.3 | 64.2 | 80.3 | 57.7 | — |
| AbstentionBenchAbstention F1↑ higher is better | 72.5best | 67.8 | 65.7 | 66.2 | 65.9 | 70.7 | — |
| Safety and refusal | |||||||
| XSTestOver-refusal on safe (%)↓ lower is better | 10.4best | 12.4 | 16.8 | 14.0 | 24.4 | 18.0 | — |
| XSTestCompliance on unsafe (%)↓ lower is better | 10.5 | 10.0 | 4.5 | 11.0 | 1.5 | 4.5 | — |
Models evaluated and why.
The comparison group covers three fronts: the closed frontier, the open-weights frontier and the Brazilian ecosystem. All were accessed through commercial APIs, on the same date.
Closed frontier
GPT-5.5 and GPT-5.4, from OpenAI, the current and the previous generation. Grok 4.6, from xAI.
Two reference labsOpen-weights frontier
DeepSeek-V4-Pro and Kimi K2.6, the two strongest Chinese open-weights models available through an API.
Published weightsBrazilian ecosystem
Sabiá 4 Thinking. Where the vendor's published protocol is comparable, its figure appears, marked with †.
Brazilian modelHow we evaluate.
- Same protocol. Every model gets the same messages, on the same date, with the same run code. Medium reasoning effort where the model accepts it. No tools and no search, except those the benchmark itself defines.
- One run per item. An unanswered question scores zero and stays in the denominator. That includes content-filter refusals, from the provider or from the API itself, counted separately.
- Automatic judge. Where the benchmark calls for a judge, we use an automatic judge with the benchmark's official prompt, without the original judge model. A different judge can label the same answer differently, and each page says so.
- Confidence interval. A 95% CI on every figure: Wilson for proportions, percentile bootstrap for F1, indices and calibration.
- Contamination. Public datasets may be in any model's training data, Genesys PI included. Each benchmark page says what that means for that result.
- Sealed data. Each dataset is pinned by commit and SHA-256. Each round's package is deterministic: the same content yields the same hash. The commands to rerun are on every page.
- Label. Every figure carries LUA Record, run by LUA Vision. It is not an independent evaluation.
LUA benchmarks.
Benchmarks LUA Vision is building for capabilities no public test measures in Portuguese. Methodology published, round in preparation. No figures yet.
| Benchmark | What it measures | Status |
|---|---|---|
| BL-NAOSEI · Knows what it doesn't know | Truthfulness, abstention and calibration in Brazilian Portuguese | Methodology published |
| BL-SAU · Health | Exam, protocol-based care, lay urgency, over-refusal | Methodology published |
| BL-KIDS · Children | Safety, age fit, correctness and tone, ages 4 to 12 | Methodology published |
Access for researchers.
Download the packages, check the hashes and run them. To rerun the LUA column, request evaluation access to the API.