LUA VISION
Bar exam

LUA would pass the Bar exam, on both stages, across the 7 most recent editions.

1st stage, objective99.3%IC 95% 98.2–99.7Would pass: yes, in all 7 editions
2nd stage, essay9.41 /10IC 95% 9.19–9.59Would pass: yes, in all 21 exams
See the full exam →
ENEM

LUA gets 99.7% of the 2024 and 2025 ENEM objective questions right.

Objective99.7%IC 95% 98.4–100.0

354 of 354 questions, 100% image coverage.

See the full exam →

LUA Record · run by LUA Vision · not an independent evaluation. No figure was rounded.

Genesys PI evaluations.

LUA Vision is a Brazilian language model lab. We evaluate Genesys PI on public benchmarks and on in-house benchmarks, side by side with frontier models, on the same date and under the same protocol. Every result carries a confidence interval and a SHA-256 sealed package so anyone can redo the math.

Results by capability.

LUA Genesys PI House first, then the comparison models. Each row leads to the benchmark page, with the confidence interval, each model's setting and the files.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

Genesys PI has the best result in the comparison on 7 of the 10 metrics. A tie on the best value counts for every tied model.

BenchmarkLUA Genesys PIGPT-5.5GPT-5.4Grok 4.6DeepSeek-V4-ProKimi K2.6Sabiá 4
Specialized knowledge in Portuguese
ENAMEDAccuracy (%)↑ higher is better96.6best96.694.395.494.393.194.4†
POSCOMPAccuracy (%)↑ higher is better95.495.495.496.293.195.490.8†
BLUEXAccuracy (%)↑ higher is better96.096.395.496.393.894.193.0†
Agents and tool use
Ticket-Bench PTAccuracy (%)↑ higher is better87.4best86.186.8—‡85.484.8—
Truthfulness and calibration
SimpleQA VerifiedF1↑ higher is better75.2best65.545.955.346.239.9—
AA-Omniscience-PublicOmniscience Index↑ higher is better56.8best35.320.737.5-2.719.2—
HalluLens · MixedEntitiesFalse acceptance (%)↓ lower is better33.0best89.267.364.280.357.7—
AbstentionBenchAbstention F1↑ higher is better72.5best67.865.766.265.970.7—
Safety and refusal
XSTestOver-refusal on safe (%)↓ lower is better10.4best12.416.814.024.418.0—
XSTestCompliance on unsafe (%)↓ lower is better10.510.04.511.01.54.5—
Truthfulness and safety: a single automatic judge per benchmark, the same for all six models, outside the evaluated model families.† Figure published by the vendor on 23/06/2026, under the vendor's protocol. Not run by LUA Vision and no confidence interval was published.‡ Environment error with page_number: every call failed. See defects.Brazil Bench run on 24/09/2026; truthfulness and safety from 25/09/2026 to 27/09/2026. LUA Genesys PI House, medium effort, public API. LUA Record · run by LUA Vision · not an independent evaluation.

Models evaluated and why.

The comparison group covers three fronts: the closed frontier, the open-weights frontier and the Brazilian ecosystem. All were accessed through commercial APIs, on the same date.

Closed frontier

GPT-5.5 and GPT-5.4, from OpenAI, the current and the previous generation. Grok 4.6, from xAI.

Two reference labs

Open-weights frontier

DeepSeek-V4-Pro and Kimi K2.6, the two strongest Chinese open-weights models available through an API.

Published weights

Brazilian ecosystem

Sabiá 4 Thinking. Where the vendor's published protocol is comparable, its figure appears, marked with †.

Brazilian model

How we evaluate.

  1. Same protocol. Every model gets the same messages, on the same date, with the same run code. Medium reasoning effort where the model accepts it. No tools and no search, except those the benchmark itself defines.
  2. One run per item. An unanswered question scores zero and stays in the denominator. That includes content-filter refusals, from the provider or from the API itself, counted separately.
  3. Automatic judge. Where the benchmark calls for a judge, we use an automatic judge with the benchmark's official prompt, without the original judge model. A different judge can label the same answer differently, and each page says so.
  4. Confidence interval. A 95% CI on every figure: Wilson for proportions, percentile bootstrap for F1, indices and calibration.
  5. Contamination. Public datasets may be in any model's training data, Genesys PI included. Each benchmark page says what that means for that result.
  6. Sealed data. Each dataset is pinned by commit and SHA-256. Each round's package is deterministic: the same content yields the same hash. The commands to rerun are on every page.
  7. Label. Every figure carries LUA Record, run by LUA Vision. It is not an independent evaluation.

LUA benchmark method →

LUA benchmarks.

Benchmarks LUA Vision is building for capabilities no public test measures in Portuguese. Methodology published, round in preparation. No figures yet.

BenchmarkWhat it measuresStatus
BL-NAOSEI · Knows what it doesn't knowTruthfulness, abstention and calibration in Brazilian PortugueseMethodology published
BL-SAU · HealthExam, protocol-based care, lay urgency, over-refusalMethodology published
BL-KIDS · ChildrenSafety, age fit, correctness and tone, ages 4 to 12Methodology published

Access for researchers.

Download the packages, check the hashes and run them. To rerun the LUA column, request evaluation access to the API.