AA-Omniscience-Public
Six hundred knowledge questions across 42 topics in six economically relevant areas. It penalizes errors and does not penalize abstention.
Results.
Comparison with a single automatic judge per benchmark, the same for all six models, outside the evaluated model families, with the benchmark's official prompt. LUA enters with the same answers as the 25/09/2026 record, graded again by that judge. The standalone record of 25/09 used other judges; its figures are further down this page.
In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).
| Model | Omniscience Index ↑ higher is better | Accuracy (%) ↑ higher is better | Abstention (%) | Hallucination rate (%) ↓ lower is better |
|---|---|---|---|---|
| LUA Genesys PI House | 56.850.8–62.7best | 69.765.9–73.2best | 16.313.6–19.5 | 42.335.4–49.6 |
| GPT-5.5 | 35.327.8–43.0 | 66.362.5–70.0 | 1.81.0–3.3 | 92.187.5–95.1 |
| GPT-5.4 | 20.713.2–28.0 | 56.852.8–60.7 | 5.53.9–7.6 | 83.878.8–87.8 |
| Grok 4.6 | 37.530.0–45.0 | 66.362.5–70.0 | 4.02.7–5.9 | 85.680.1–89.8 |
| DeepSeek-V4-Pro | -2.7-10.2–5.2 | 46.742.7–50.7 | 3.52.3–5.3 | 92.589.1–94.9 |
| Kimi K2.6 | 19.213.2–25.2 | 39.235.3–43.1 | 40.236.3–44.1 | 32.928.3–37.9 |
| Sabiá 4 Thinking | — | — | — | — |
Omniscience Index · higher is better
Axis from -25 to 75. Black bar: 95% confidence interval.
External references
Published by third parties (artificialanalysis.ai, artificialanalysis.ai, artificialanalysis.ai, artificialanalysis.ai, artificialanalysis.ai, artificialanalysis.ai, artificialanalysis.ai, accessed 26/09/2026): GPT-5.5 (xhigh), Omniscience Index (escala -100 a 100) 20.5; GPT-5.5 (xhigh), Acurácia (%) 58.0; GPT-5.5 (xhigh), Taxa de alucinação (%) 89.0; GPT-5.4 (xhigh), Omniscience Index (escala -100 a 100) 5.8; GPT-5.4 (xhigh), Acurácia (%) 50.8; GPT-5.4 (xhigh), Taxa de alucinação (%) 91.7; Grok 4.6 (high), Omniscience Index (escala -100 a 100) 30.5; Grok 4.6 (high), Acurácia (%) 48.2; Grok 4.6 (high), Taxa de alucinação (%) 34.3; DeepSeek V4 Pro 0813 (Reasoning, Max Effort), Omniscience Index (escala -100 a 100) 0.8; DeepSeek V4 Pro 0813 (Reasoning, Max Effort), Acurácia (%) 49.1; DeepSeek V4 Pro 0813 (Reasoning, Max Effort), Taxa de alucinação (%) 94.8; DeepSeek V4 Pro 0424 (Reasoning, Max Effort), Omniscience Index (escala -100 a 100) -10.7; DeepSeek V4 Pro 0424 (Reasoning, Max Effort), Acurácia (%) 43.0; DeepSeek V4 Pro 0424 (Reasoning, Max Effort), Taxa de alucinação (%) 94.1; Kimi K2.6, Omniscience Index (escala -100 a 100) 5.3; Kimi K2.6, Acurácia (%) 32.6; Kimi K2.6, Taxa de alucinação (%) 40.5; GPT-5.5 (xhigh), Acurácia (%) no lançamento 57.0; GPT-5.5 (xhigh), Taxa de alucinação (%) no lançamento 86.0. Judge and setting differ from this page. These are official-index figures, on 6,000 questions, not the public subset of 600 used here.
Comparison files
| Per-item result | SHA-256 | Download |
|---|---|---|
resultados/omniscience.house.rejulgado.jsonl | 083505b520d6fe6dc77a57600f7855d1b7c30a42a836ab38792fbcb0e1d675c9 | withheld: names the judge |
resultados/omniscience.gpt55.jsonl | cd661e721d0e8d784dda69eb87c0a37c189b668d0406c05aa8cbbdd56a19bf35 | withheld: names the judge |
resultados/omniscience.gpt54.jsonl | c19dee5b83b271567bac61c9e4de8e08726a1bc739444493e85edf8f43e22cd9 | withheld: names the judge |
resultados/omniscience.grok.jsonl | ce46be86a95fb48b2e142bdb7fc7c0e063e3ec3888ec0b534ad0a53ea03a636c | withheld: names the judge |
resultados/omniscience.deepseek.jsonl | 0ff781c708e5c227bf867b060564c31aa508944e26d7be338daf8d4ad5d1721b | withheld: names the judge |
resultados/omniscience.kimi.jsonl | c67bf6078226b576862c5afbc32ca440a2b11d40b5d89f6ceb553be6f016eac8 | withheld: names the judge |
How the comparison was run.
Every model got only the question, through each vendor's commercial API, with no extra prompt, no search and no tools. One run per item. Unanswered questions stay in the denominator.
Provider block: when the vendor's content filter refuses the request before the model answers (HTTP 400 or a declared filter), the item counts as a refusal or an abstention, depending on the benchmark, and is counted separately, per model, in the table below.
| Model | Access | Reasoning effort | Provider blocks | No answer |
|---|---|---|---|---|
| LUA Genesys PI House | public API api.lua.vision | medium | 0 | 0 |
| GPT-5.5 | API comercial | medium | 0 | 0 |
| GPT-5.4 | API comercial | medium | 0 | 0 |
| Grok 4.6 | API comercial | model default | 1 | 1 |
| DeepSeek-V4-Pro | API comercial | model default | 1 | 0 |
| Kimi K2.6 | API comercial | model default | 1 | 3 |
Standalone record, 25/09/2026.
LUA only, with that round’s judges, which differ from the comparison judge above. These figures do not compare with the table above.
| Metric | Value | 95% CI |
|---|---|---|
| Omniscience Index | 54.5 | 48.7–60.3 |
| Accuracy (%) | 68.7 | 64.8–72.2 |
| Abstention (%) | 15.8 | 13.1–19.0 |
| Hallucination rate (%) | 45.2 | 38.3–52.4 |
Omniscience Index = 100 × (correct − wrong) / total. Hallucination rate = wrong / (partial + wrong + abstention): how often LUA guesses wrong when it does not know. n = 600: 412 correct, 85 wrong, 8 partial, 95 abstentions. Abstention carries a Wilson interval computed from those counts.
A hallucination rate always goes out next to accuracy and abstention. A model that refuses everything scores zero hallucination, and the table shows it.
How it was run.
| Call | LUA Genesys PI House, genesys-pi-house, public API api.lua.vision, reasoning effort medium. No temperature, no token cap, no extra prompt, no search and no tools. |
|---|---|
| Dataset | ArtificialAnalysis/AA-Omniscience-Public, revision e4883edb, CC BY 4.0. arXiv 2511.13029 |
| Items | 600 of 600. |
| Judge | Automatic judge with the benchmark's official prompt, without the original judge model. The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down. |
| Runs | One per item. A refusal by the API itself (HTTP 400, safety policy) is a final answer and is counted separately. |
| Interval | Wilson for proportions. Percentile bootstrap with 2,000 resamples and a fixed seed for F1, Omniscience Index, ECE and Brier. |
Package and files.
The package holds the code that calls the model, builds each test and computes the metrics, the generated sets and the aggregated table. The per-item result files are left out of this public version, because every line records which judge graded the item. The SHA-256 of each one is below, matching the round's table.
ff68d8db16a4d3944e635b50a4a9b28f7009fb0cb94fa37510c1e579681b0f82| File | Size | Download |
|---|---|---|
registro-veracidade-2026-09.tar.gz | 41 KB | Download |
| Per-item result | SHA-256 | Download |
|---|---|---|
resultados/omniscience.house.jsonl | cb3a6155906961f51ef866c53b7beeec34b287b8414681ff36457d07a8419b0e | withheld: names the judge |
Left out of this public version for the same reason: modelos.mjs, benches/abstention.mjs, benches/hallulens.mjs, benches/omniscience.mjs, benches/orbench.mjs, benches/orbench_hard_j2.mjs, benches/orbench_toxic_j2.mjs, benches/simpleqa.mjs, benches/simpleqa_conf.mjs, benches/xstest.mjs, README.md, colunas.json, tabela.json. Without them the package does not run on its own. The full version depends on a pending decision about naming the judge.
Limitations.
- This is the public subset, not the maintainer's official index, which uses the full set.
- The judge that graded the answers is not the one the benchmark authors used. The grading prompt is the official one, unchanged, but a different judge can label the same answer differently. That can move the figures up or down.