LUA VISION

AbstentionBench

Whether the model abstains when it should: questions with no known answer, underspecified, with a false premise, subjective or out of date.

RunLUA Record · run by LUA Vision · not an independent evaluation

Results.

Comparison with a single automatic judge per benchmark, the same for all six models, outside the evaluated model families, with the benchmark's official prompt. LUA enters with the same answers as the 25/09/2026 record, graded again by that judge. The standalone record of 25/09 used other judges; its figures are further down this page.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

ModelAbstention F1
↑ higher is better
Recall
↑ higher is better
Precision
↑ higher is better
LUA Genesys PI House72.569.8–75.0best64.961.7–68.0best82.179.0–84.8
GPT-5.567.864.8–70.654.651.3–57.989.386.4–91.7
GPT-5.465.762.7–68.653.550.1–56.885.382.0–88.0
Grok 4.666.263.4–69.152.248.9–55.690.487.5–92.7
DeepSeek-V4-Pro65.963.1–68.754.451.1–57.883.380.0–86.2
Kimi K2.670.768.1–73.360.457.1–63.685.182.1–87.8
Sabiá 4 Thinking———

Abstention F1 · higher is better

LUA Genesys PI House
best result72.5IC 69.8–75.0
Kimi K2.6
70.7IC 68.1–73.3
GPT-5.5
67.8IC 64.8–70.6
Grok 4.6
66.2IC 63.4–69.1
DeepSeek-V4-Pro
65.9IC 63.1–68.7
GPT-5.4
65.7IC 62.7–68.6

Axis from 60 to 80. Black bar: 95% confidence interval.

The small line in each cell is the 95% CI.Items graded, where not all: DeepSeek-V4-Pro: 1999 of 2000; Grok 4.6: 1997 of 2000; Kimi K2.6: 1999 of 2000.Sabiá 4 appears only where the vendor's published protocol is comparable: multiple choice with exact-match accuracy.

Comparison files

Per-item resultSHA-256Download
resultados/abstention.house.rejulgado.jsonl0230d79a74260a4aced7d29571850a79970fe73bd2326485f6497231a099bbc2withheld: names the judge
resultados/abstention.gpt55.jsonl7c88a9f9e4875bc3cff1d4ecb15d49f481dfc2894746d54da47e70e08eef16d7withheld: names the judge
resultados/abstention.gpt54.jsonl8de18028f01b92631ee3488f0728fb18f8d87ec5c27e714563b0353d0d0cc6ffwithheld: names the judge
resultados/abstention.grok.jsonl6bb01490b6656a6ff07a98d0a5d81df4f19c3ea744cc82c107fa6f622fa781d5withheld: names the judge
resultados/abstention.deepseek.jsonldceef4950539caa8a63705511bb0fd4547f27006409994fb18200ff058ffb0bbwithheld: names the judge
resultados/abstention.kimi.jsonl38b2993bdd4919d40e310ccd8797669fad3a5f4d3e0e6bf595ba7aafca8b4afcwithheld: names the judge

How the comparison was run.

Every model got only the question, through each vendor's commercial API, with no extra prompt, no search and no tools. One run per item. Unanswered questions stay in the denominator.

Provider block: when the vendor's content filter refuses the request before the model answers (HTTP 400 or a declared filter), the item counts as a refusal or an abstention, depending on the benchmark, and is counted separately, per model, in the table below.

ModelAccessReasoning effortProvider blocksNo answer
LUA Genesys PI Housepublic API api.lua.visionmedium140
GPT-5.5API comercialmedium120
GPT-5.4API comercialmedium130
Grok 4.6API comercialmodel default142
DeepSeek-V4-ProAPI comercialmodel default221
Kimi K2.6API comercialmodel default241

Standalone record, 25/09/2026.

LUA only, with that round’s judges, which differ from the comparison judge above. These figures do not compare with the table above.

MetricValue95% CI
Recall66.563.3–69.6
Precision75.772.5–78.6
F170.868.3–73.4

n = 2,000. 855 items call for abstention; LUA abstained on 752, and 569 of those abstentions were right. Recall and precision go together: abstaining on everything gives high recall and low precision. 14 items were refused by the API itself and count as abstention.

By question type

TypenCall for abstentionRecall (%)Precision (%)
underspecified context1,20751462.6 IC 58.4–66.784.7 IC 80.8–88.0
false premise22112190.1 IC 83.5–94.282.6 IC 75.2–88.1
unknown answer1516168.9 IC 56.4–79.167.7 IC 55.4–78.0
subjective1276826.5 IC 17.4–38.090.0 IC 69.9–97.2
underspecified intent1202295.5 IC 78.2–99.236.2 IC 25.1–49.1
stale1042572.0 IC 52.4–85.731.0 IC 20.6–43.8
safety504090.0 IC 76.9–96.097.3 IC 86.2–99.5
no scenario in the paper20475.0 IC 30.1–95.460.0 IC 23.1–88.2

A type with few items calling for abstention has a wide interval. Read the interval before the figure.

How it was run.

CallLUA Genesys PI House, genesys-pi-house, public API api.lua.vision, reasoning effort medium. No temperature, no token cap, no extra prompt, no search and no tools.
DatasetAbstentionBench, fast subset, commit e2918417: 20 datasets × 100 items. arXiv 2506.09038
Items2,000. The fast index lists 21 datasets. Averitec was left out because it has no loader in the benchmark repository.
JudgeAutomatic judge with the benchmark's official prompt, without the original judge model. The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.
RunsOne per item. A refusal by the API itself (HTTP 400, safety policy) is a final answer and is counted separately.
IntervalWilson for proportions. Percentile bootstrap with 2,000 resamples and a fixed seed for F1, Omniscience Index, ECE and Brier.

Package and files.

The package holds the code that calls the model, builds each test and computes the metrics, the generated sets and the aggregated table. The per-item result files are left out of this public version, because every line records which judge graded the item. The SHA-256 of each one is below, matching the round's table.

registro-veracidade-2026-09.tar.gzff68d8db16a4d3944e635b50a4a9b28f7009fb0cb94fa37510c1e579681b0f82
FileSizeDownload
registro-veracidade-2026-09.tar.gz41 KBDownload
Per-item resultSHA-256Download
resultados/abstention.house.jsonld6ea6160050360e74d906e6163e71c9b4aa2b658383cd931b15d7b37d4f78ba5not in this version

Left out of this public version for the same reason: modelos.mjs, benches/abstention.mjs, benches/hallulens.mjs, benches/omniscience.mjs, benches/orbench.mjs, benches/orbench_hard_j2.mjs, benches/orbench_toxic_j2.mjs, benches/simpleqa.mjs, benches/simpleqa_conf.mjs, benches/xstest.mjs, README.md, colunas.json, tabela.json. Without them the package does not run on its own. The full version depends on a pending decision about naming the judge.

Limitations.

  • Recall and precision go out together: abstaining on everything gives high recall and low precision.
  • The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.