LUA VISION

SimpleQA Verified

A thousand short factual questions with checkable answers and no search. The model can be right, wrong or decline. It measures whether it guesses when it does not know.

RunLUA Record · run by LUA Vision · not an independent evaluation

Results.

Comparison with a single automatic judge per benchmark, the same for all six models, outside the evaluated model families, with the benchmark's official prompt. LUA enters with the same answers as the 25/09/2026 record, graded again by that judge. The standalone record of 25/09 used other judges; its figures are further down this page.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

ModelF1
↑ higher is better
Accuracy (%)
↑ higher is better
Not attempted (%)Accuracy when attempted (%)
↑ higher is better
ECE ×100 (calibration error)
↓ lower is better
Brier
↓ lower is better
LUA Genesys PI House75.272.0–78.0best72.169.2–74.8best8.36.7–10.278.675.9–81.2best12.410.3–14.80.1530.140–0.170best
GPT-5.565.563.0–68.065.462.4–68.30.30.1–0.965.662.6–68.516.814.1–19.30.2020.180–0.220
GPT-5.445.943.0–49.045.542.4–48.61.61.0–2.646.243.1–49.414.512.0–17.20.1970.180–0.210
Grok 4.655.352.0–58.054.951.8–58.01.61.0–2.655.852.7–58.910.99.0–14.10.2130.200–0.230
DeepSeek-V4-Pro46.243.0–49.044.841.7–47.96.14.8–7.847.744.5–50.933.430.5–36.40.3300.310–0.350
Kimi K2.639.937.0–43.035.532.6–38.522.219.7–24.945.642.2–49.111.69.3–14.40.1880.170–0.200
Sabiá 4 Thinking——————

F1 · higher is better

LUA Genesys PI House
best result75.2IC 72.0–78.0
GPT-5.5
65.5IC 63.0–68.0
Grok 4.6
55.3IC 52.0–58.0
DeepSeek-V4-Pro
46.2IC 43.0–49.0
GPT-5.4
45.9IC 43.0–49.0
Kimi K2.6
39.9IC 37.0–43.0

Axis from 30 to 85. Black bar: 95% confidence interval.

The small line in each cell is the 95% CI.Not attempted and abstention have no better direction: they are there to read accuracy and hallucination together.Sabiá 4 appears only where the vendor's published protocol is comparable: multiple choice with exact-match accuracy.

External references

Published by third parties (www.kaggle.com, accessed 26/09/2026): GPT-5.4, Score (%) 30.5. Judge and setting differ from this page.

Comparison files

Per-item resultSHA-256Download
resultados/simpleqa.house.jsonl4bf96842d6d64026ea3b62f35d0158056af4dc4fe0fa3e36d411922f7cf48492withheld: names the judge
resultados/simpleqa.gpt55.jsonl715d4903057616edcc49a722f132b0f6c94646fb295cbf67487db85f7a148967withheld: names the judge
resultados/simpleqa.gpt54.jsonl040db462820c1b17dd87e54d589a53cb4ec57195661a60e7726bb1ad02f7ee11withheld: names the judge
resultados/simpleqa.grok.jsonl3b5bc95969f86c868755b5abd1f672ac3e863af2bd13d523bc74d6d6206d17f0withheld: names the judge
resultados/simpleqa.deepseek.jsonlcbc62179b0b54a9854d1b55f79a81f4422a84423639845c2baa6377a7a1aa0e0withheld: names the judge
resultados/simpleqa.kimi.jsonld5de43ae0dc1fd27103c64a0b526948d05b7515f78a9aa39d406c1f27312bd68withheld: names the judge
resultados/simpleqa_conf.house.rejulgado.jsonle5c399079556bcb992b5356111cd41f0018c6531e5a19e9f967c8768f964a987withheld: names the judge
resultados/simpleqa_conf.gpt55.jsonle12bc360df7519ce9c8f678ca24ca3ca1a8225c773dca9907a3005c29738e01awithheld: names the judge
resultados/simpleqa_conf.gpt54.jsonl9550cc0c2e5e18f98676b1c26fc07dac96f7bf0dc82f4a95bb132fc8a92bfe85withheld: names the judge
resultados/simpleqa_conf.grok.jsonlb6796a56d4d686df5561570423b252893061eacd94f14249670684e8e52a3d9cwithheld: names the judge
resultados/simpleqa_conf.deepseek.jsonld4324f242ecb44a6e32a69fe02afb949e8b1d56fbc1f050cf6e34467673d51cdwithheld: names the judge
resultados/simpleqa_conf.kimi.jsonlb85a95a3cb967006dd5423ba65ebf4f2387fea11922c1b21ee3675c27db2468fwithheld: names the judge

How the comparison was run.

Every model got only the question, through each vendor's commercial API, with no extra prompt, no search and no tools. One run per item. Unanswered questions stay in the denominator.

Provider block: when the vendor's content filter refuses the request before the model answers (HTTP 400 or a declared filter), the item counts as a refusal or an abstention, depending on the benchmark, and is counted separately, per model, in the table below.

ModelAccessReasoning effortProvider blocksNo answer
LUA Genesys PI Housepublic API api.lua.visionmedium00
GPT-5.5API comercialmedium00
GPT-5.4API comercialmedium00
Grok 4.6API comercialmodel default00
DeepSeek-V4-ProAPI comercialmodel default20
Kimi K2.6API comercialmodel default60

Standalone record, 25/09/2026.

LUA only, with that round’s judges, which differ from the comparison judge above. These figures do not compare with the table above.

MetricValue95% CI
F175.272.0–78.0
Accuracy (%)72.169.2–74.8
Not attempted (%)8.36.7–10.2
Accuracy when attempted (%)78.675.9–81.2
Error when attempted (%)21.418.8–24.1

n = 1,000: 721 correct, 196 wrong, 83 not attempted. Error when attempted is the complement of accuracy when attempted, with the same interval mirrored. Not attempted always sits next to accuracy: a model that never tries scores zero error.

Calibration.

A second pass with the SimpleQA paper's prompt that asks for the answer and a confidence from 0 to 100. LUA protocol on the SimpleQA Verified dataset. This pass asks for a guess, so its accuracy does not replace the first pass's.

Finding: LUA states more confidence than it earns. In the confidence pass, mean confidence was 85.5 and accuracy was 73.4%. On the 692 answers with confidence 90 to 100, it was right 88.9% of the time. Below 80 confidence, accuracy sits between 6% and 35%, well under what it states.

MetricValue95% CI
ECE, 10 bins (points)12.6710.64–15.11
Brier0.1530.130–0.170
Accuracy in this pass (%)73.470.6–76.0
Mean confidence85.5—
Reliability diagramAccuracy by stated-confidence bin, with the perfect-calibration diagonal. The values are in the table below.002020404060608080100100n=18n=14n=33n=17n=13n=22n=29n=33n=129n=692Stated confidence (%)Accuracy (%)

bin accuracybin mean confidenceperfect calibration

BinnMean confidenceAccuracy (%)
0-10183.35.6
10-201413.628.6
20-303322.69.1
30-401734.635.3
40-501343.030.8
50-602255.531.8
60-702963.624.1
70-803375.233.3
80-9012985.558.9
90-10069296.688.9

Diagram drawn from the confidence pass result file (last graded line per item), checked against the round table.

How it was run.

CallLUA Genesys PI House, genesys-pi-house, public API api.lua.vision, reasoning effort medium. No temperature, no token cap, no extra prompt, no search and no tools.
DatasetSimpleQA Verified, 1,000 questions, revision 0dc97e0d. arXiv 2509.07968
Items1,000 of 1,000.
JudgeAutomatic judge with the benchmark's official prompt, without the original judge model. The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.
RunsOne per item. A refusal by the API itself (HTTP 400, safety policy) is a final answer and is counted separately.
IntervalWilson for proportions. Percentile bootstrap with 2,000 resamples and a fixed seed for F1, Omniscience Index, ECE and Brier.

Package and files.

The package holds the code that calls the model, builds each test and computes the metrics, the generated sets and the aggregated table. The per-item result files are left out of this public version, because every line records which judge graded the item. The SHA-256 of each one is below, matching the round's table.

registro-veracidade-2026-09.tar.gzff68d8db16a4d3944e635b50a4a9b28f7009fb0cb94fa37510c1e579681b0f82
FileSizeDownload
registro-veracidade-2026-09.tar.gz41 KBDownload
Per-item resultSHA-256Download
resultados/simpleqa.house.jsonl4bf96842d6d64026ea3b62f35d0158056af4dc4fe0fa3e36d411922f7cf48492withheld: names the judge
resultados/simpleqa_conf.house.jsonldfe064c29b9fa9fec347a38365a3fa9e8c7a24606c84cd3f5557e3b1d1fd2969withheld: names the judge

Left out of this public version for the same reason: modelos.mjs, benches/abstention.mjs, benches/hallulens.mjs, benches/omniscience.mjs, benches/orbench.mjs, benches/orbench_hard_j2.mjs, benches/orbench_toxic_j2.mjs, benches/simpleqa.mjs, benches/simpleqa_conf.mjs, benches/xstest.mjs, README.md, colunas.json, tabela.json. Without them the package does not run on its own. The full version depends on a pending decision about naming the judge.

Limitations.

  • Questions in English. It measures memory, not search.
  • The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.