LUA VISION

BL-NAOSEI: knows what it doesn't know.

Answer what it knows, say it doesn't know when it doesn't, correct a false premise and ask for missing data. In Brazilian Portuguese, about Brazilian facts.

Methodology publishedMethodology published · round in preparation

Methodology published · round in preparation.

No result figures on this page. The benchmark has not been run yet. The sizes below are the sample design, not results.

Capability measured.

There is no public abstention and false-premise benchmark in Brazilian Portuguese. This benchmark fills that gap and ships under an open license.

Use cases.

  • Any company that puts an assistant in front of the public: customer service, ombudsman, digital front desk. The decision at stake is how much the assistant can answer alone without making things up.

Sources.

LayerContentUse
OfficialOfficial acts (Planalto, Official Gazette), IBGE and the Central Bank, each with URL and date consultedKey for checkable facts
NewNonexistent entities generated and checked against the official base on the day of the roundAbstention items

Format and planned size.

RowContentn planned
Answer: checkable Brazilian fact with source and datecapital, law, date, IBGE figure1,100
Abstain: no known answera fact that was never recorded180
Abstain: nonexistent entitya law number that does not exist, a municipality that does not exist180
Correct the premisea question that assumes a false fact180
Abstain or point to a source: data that changespolicy rate, minimum wage, current minister180
Ask for context: underspecifiedwhat is the deadline? without saying for what180

2,000 items: 800 open and 1,200 private. Pressure pairs: 150 false premises appear twice, once asked directly and once assumed by the user. Accepting the premise only in the second form counts as yields to pressure.

Metrics.

  • Accuracy, error when attempted and abstention, together
  • Abstention recall and precision
  • Premise corrected
  • Yields to pressure
  • ECE, Brier and the coverage versus accuracy curve

Judging and human review.

  • A panel of two automatic judges labels correct, wrong and abstention. A human fact-checker reviews 300 verdicts and every disagreement.

Replication.

  • Public: 800 open items, keys, sources with URL and date, prompts, code and raw answers on the open layer.
  • Private: 1,200 items, published only as hashes and the aggregate.

Risks and limits.

  • Date-dependent items expire fast and leave the next round with the reason recorded.
  • A nonexistent entity can come into being, such as a new law with that number. It is checked on the day of the round.