LUA VISION

Ticket-Bench in Portuguese: the agent buys the ticket.

An agent takes a fan's request, looks up matches and prices through tools in a simulated environment and buys the ticket. It counts as correct when the final purchase matches the expected one, in one attempt (pass@1).

LUA Record · run by LUA Vision · not an independent evaluation

Results.

In each row, the best value is highlighted. Mind the direction: for some metrics, lower is better (hallucination, over-refusal, calibration error).

ModelAccuracy (%)
↑ higher is better
95% CIAnsweredRefused or blockedNo answer
LUA Genesys PI House87.4best76.0–96.815100
GPT-5.586.173.9–96.515100
GPT-5.486.874.8–96.615100
Grok 4.6‡——390112
DeepSeek-V4-Pro85.474.5–95.115100
Kimi K2.684.873.0–95.315100

Accuracy (%) · higher is better

LUA Genesys PI House
best result87.4IC 76.0–96.8
GPT-5.4
86.8IC 74.8–96.6
GPT-5.5
86.1IC 73.9–96.5
DeepSeek-V4-Pro
85.4IC 74.5–95.1
Kimi K2.6
84.8IC 73.0–95.3
Grok 4.6 ‡
—

Axis from 65 to 100. Black bar: 95% confidence interval.

‡ Environment error with page_number: every call failed. See defects.n = 151. 95% interval by percentile bootstrap, 4,000 resamples, fixed seed. Variations of the same question template are resampled together, because they are not independent.Sabiá 4 appears only where the vendor's published protocol is comparable: multiple choice with exact-match accuracy.

The errors in the public Portuguese dataset, described below, weigh on every model the same way. The Ticket-Bench paper reports GPT-5 at 0.87 in Portuguese, the same level as this round. Compare the models in this table with one another.

How it was run.

Date24/09/2026
DatasetTropicAI-Research/Ticket-Bench, commit 7318bc718a07
Items151. None excluded.
ScoringFinal purchase equal to the expected one (pass@1). No judge.
RunsOne per item, pass@1.
FailuresAn unanswered question scores zero and stays in the denominator. That includes refusals by the provider's content filter and refusals by the model itself.
Interval95% interval by percentile bootstrap, 4,000 resamples, fixed seed. Variations of the same question template are resampled together, because they are not independent.

Setting per model.

Every model got the same messages, through the commercial API of each vendor or of the cloud where the model is published, on the same date. No extra system prompt, no temperature, no tools beyond those the benchmark defines.

ModelAccessReasoning effort
LUA Genesys PI Housepublic API api.lua.visionmedium
GPT-5.5API comercialmedium
GPT-5.4API comercialmedium
Grok 4.6API comercialmodel default
DeepSeek-V4-ProAPI comercialmodel default
Kimi K2.6API comercialmodel default

Protocol decisions.

  • The original code sets temperature 0.7. Reasoning models reject that parameter, so no column gets a temperature.
  • State, tools, system prompt, the 20-iteration limit and the scoring routine are the authors' own, imported straight from the clone.

Defects found in the data.

  • Portuguese data: question template 12 was translated as "next match" where the original asks for the "cheapest", and the answer keys for questions 6, 11 and 17 follow neither the question nor the repository's own Portuguese standings table.
  • Environment: the listar_partidas tool declares page_number as number, but the authors' code uses the value to slice a list and needs an integer. A model that sends 1.0, valid under the schema, gets an error on every call and runs out of its 20 iterations. This happened only with Grok 4.6, which leaves this table as not comparable. The other columns send integers and never hit the limit.

Contamination.

The items are in public datasets. Any column, LUA included, may have seen these questions in training. This round did not measure contamination, so read the figure as a ceiling on what the model knows about the subject, not as proof of reasoning on unseen items.

Limitations.

  • One run per item. Run-to-run variation of the same model was not measured.
  • A single zero-shot multiple-choice prompt. Another prompt may change the order of the columns.
  • The figure holds for the model served by the API on 24/09/2026. A new version needs a new round.
  • Questions with images were left out for every column. The result says nothing about reading figures.

Reproduce.

The package holds the code that calls the models, builds each test, scores it and builds the table, plus the aggregated table for this round. The data is downloaded by preparar_dados.py at the exact versions and checked by SHA-256. If one byte changes, the script stops.

With Python 3 and Node 20 or newer:

tar -xzf bancada-brasil-2026-09.tar.gz && cd bancada-brasil-2026-09
shasum -a 256 -c <(sed -n '/^Arquivos deste pacote/,/^$/p' VERIFICAR.txt | grep -E '^[0-9a-f]{64}')
pip install huggingface_hub pyarrow
python3 preparar_dados.py          # baixa, normaliza e confere SHA-256
export LUA_API_KEY=...             # chave de avaliação da API LUA
python3 ticket_run.py --modelo house
node tabela.mjs
bancada-brasil-2026-09.tar.gz1144df20af11b5c0141614c256fa1733c698bf829be3303842839ec149a1f81111 KB · Raw per-item answers are not in this package.
FileSizeDownload
bancada-brasil-2026-09.tar.gz11 KBDownload
resultados-agregados.json7 KBDownload
VERIFICAR.txt1 KBDownload

History.

24/09/2026 · first round, published on this page.