Evaluation run by LUA with the LiveBench evaluation code, on the public questions of the 2026-01-08 set: 682 questions across five categories (data analysis, instruction following, language, math and reasoning), at temperature zero. Coding was left out because the set has no public questions in that category. Model evaluated: LUA Genesys, second generation.
42b6cb4eef86cc7b4124c300571dba9fdc46baac52530790235a679b4aa42381Sealed file: lua-genesys-god-nim-livebench-submission.tar.gz. File available on request to verify the hash.