LUA VISION

HalluLens MixedEntities

Questions about animals, plants and bacteria that do not exist, NonExistentRefusal task. It measures how often the model accepts and describes something made up.

RunLUA Record · run by LUA Vision · not an independent evaluation

Results.

Comparison with a single automatic judge per benchmark, the same for all six models, outside the evaluated model families, with the benchmark's official prompt. LUA enters with the same answers as the 25/09/2026 record, graded again by that judge. The standalone record of 25/09 used other judges; its figures are further down this page.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

ModelFalse acceptance of invented entities (%)
↓ lower is better
LUA Genesys PI House33.029.4–36.9best
GPT-5.589.286.4–91.4
GPT-5.467.363.5–71.0
Grok 4.664.260.3–67.9
DeepSeek-V4-Pro80.376.9–83.3
Kimi K2.657.753.7–61.6
Sabiá 4 Thinking—

False acceptance of invented entities (%) · lower is better

LUA Genesys PI House
best result33.0IC 29.4–36.9
Kimi K2.6
57.7IC 53.7–61.6
Grok 4.6
64.2IC 60.3–67.9
GPT-5.4
67.3IC 63.5–71.0
DeepSeek-V4-Pro
80.3IC 76.9–83.3
GPT-5.5
89.2IC 86.4–91.4

Axis from 20 to 100. Black bar: 95% confidence interval.

The small line in each cell is the 95% CI.Sabiá 4 appears only where the vendor's published protocol is comparable: multiple choice with exact-match accuracy.

Comparison files

Per-item resultSHA-256Download
resultados/hallulens.house.rejulgado.jsonl6f7890d3939ad22423bd51f9e89ff54e02d4f641e3bf8da691c0ea01ff08b2abwithheld: names the judge
resultados/hallulens.gpt55.jsonl31174c8f7c0343ab16cf55ed2a42450dc0fe5210826ddad6cd345dd1e3b597c3withheld: names the judge
resultados/hallulens.gpt54.jsonl3cd3c0ec7de9881a56f473957b5425198cfffda237f26957e84fe8e9ffa9d2d8withheld: names the judge
resultados/hallulens.grok.jsonl3d92d38fbe2f7f08c056edc1833cc32d67cb8daf4c582e1b1357036cb05df8fdwithheld: names the judge
resultados/hallulens.deepseek.jsonlcd209cb0cc726b12dececfb8049992e18ee5a447f30f0b5f5d6001114f8fac06withheld: names the judge
resultados/hallulens.kimi.jsonl34b386c4e130b7888718eab9e1cd9189a94bb66c3b55aaf5b806219aa5b6f479withheld: names the judge

How the comparison was run.

Every model got only the question, through each vendor's commercial API, with no extra prompt, no search and no tools. One run per item. Unanswered questions stay in the denominator.

Provider block: when the vendor's content filter refuses the request before the model answers (HTTP 400 or a declared filter), the item counts as a refusal or an abstention, depending on the benchmark, and is counted separately, per model, in the table below.

ModelAccessReasoning effortProvider blocksNo answer
LUA Genesys PI Housepublic API api.lua.visionmedium00
GPT-5.5API comercialmedium00
GPT-5.4API comercialmedium00
Grok 4.6API comercialmodel default00
DeepSeek-V4-ProAPI comercialmodel default10
Kimi K2.6API comercialmodel default00

Standalone record, 25/09/2026.

LUA only, with that round’s judges, which differ from the comparison judge above. These figures do not compare with the table above.

DomainnAcceptedFalse acceptance (%)95% CI
All domains60019532.528.9–36.3
Animal20010854.047.1–60.8
Plant2005527.521.8–34.1
Bacteria2003216.011.6–21.7

False acceptance: LUA described as real an animal, plant or bacterium that does not exist. Lower is better. The errors cluster in animals.

How it was run.

CallLUA Genesys PI House, genesys-pi-house, public API api.lua.vision, reasoning effort medium. No temperature, no token cap, no extra prompt, no search and no tools.
DatasetHalluLens, NonExistentRefusal task, MixedEntities, commit 80307ac6, official generator with seed 0 and 200 names per domain. arXiv 2504.17550
Items600. The official script uses 10 names per domain; here there are 200, so the interval is useful. Answers are not cut at 256 tokens as in the official inference. The medicines domain was left out: its source needs a Kaggle account. GeneratedEntities was left out: it needs a paid search API.
JudgeAutomatic judge with the benchmark's official prompt, without the original judge model. The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.
RunsOne per item. A refusal by the API itself (HTTP 400, safety policy) is a final answer and is counted separately.
IntervalWilson for proportions. Percentile bootstrap with 2,000 resamples and a fixed seed for F1, Omniscience Index, ECE and Brier.

Package and files.

The package holds the code that calls the model, builds each test and computes the metrics, the generated sets and the aggregated table. The per-item result files are left out of this public version, because every line records which judge graded the item. The SHA-256 of each one is below, matching the round's table.

registro-veracidade-2026-09.tar.gzff68d8db16a4d3944e635b50a4a9b28f7009fb0cb94fa37510c1e579681b0f82
FileSizeDownload
registro-veracidade-2026-09.tar.gz41 KBDownload
hallulens_mixed_semente0_n200.jsonl · SHA-256 c2c11d849a74f550eaa14b3f06de0b6f855cc796fb566b4263c178544a34a5d590 KBDownload
hallulens_fonte.json · SHA-256 2091130d9bd8e4b2e69ac9e1e5d0305fe11d7c2215ce2e3cf10805689314d2951 KBDownload
Per-item resultSHA-256Download
resultados/hallulens.house.jsonl0ace3bda4aef2f948b5bebb878d5452913dd68055aa5cd7fece45f23869325cawithheld: names the judge

Left out of this public version for the same reason: modelos.mjs, benches/abstention.mjs, benches/hallulens.mjs, benches/omniscience.mjs, benches/orbench.mjs, benches/orbench_hard_j2.mjs, benches/orbench_toxic_j2.mjs, benches/simpleqa.mjs, benches/simpleqa_conf.mjs, benches/xstest.mjs, README.md, colunas.json, tabela.json. Without them the package does not run on its own. The full version depends on a pending decision about naming the judge.

Limitations.

  • Dynamic set: two rounds do not use the same questions. Comparisons hold within the same seed.
  • The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.