About LiveBench
The questions in Quick and Deep come from LiveBench, a public benchmark for AI language models. HIBN is not affiliated with LiveBench and is not endorsed by it. A HIBN score is not a LiveBench score.
What LiveBench is
LiveBench describes itself as “a benchmark for LLMs designed with test set contamination and objective evaluation in mind”. LLMs are the large language models behind AI chat apps.
It was created by Colin White and 17 co-authors. They are from Abacus.AI, New York University, Nvidia, the University of Maryland, the University of Southern California and Columbia University. LiveBench’s datasheet says: “This benchmark was funded by Abacus.AI.”
The paper is “LiveBench: A Challenging, Contamination-Limited LLM Benchmark”. LiveBench’s README says: “LiveBench appeared as a Spotlight Paper in ICLR 2025.”
LiveBench publishes a leaderboard of AI models. It adds new questions regularly, because questions that stay public can end up in the data that AI models are trained on.
LiveBench website · LiveBench on GitHub · The LiveBench paper
HIBN read these LiveBench sources on 2 October 2026: the paper (version 2), the README, the datasheet, the licence file and the maths scoring code.
What HIBN uses
HIBN uses 16 LiveBench questions of 3 types. It took them from LiveBench’s public datasets on Hugging Face.
| HIBN’s name | LiveBench’s name and dataset | In Quick | In Deep | LiveBench release |
|---|---|---|---|---|
| Logic puzzle | zebra_puzzle, in livebench/reasoning | 2 | 3 | 25 November 2024 |
| Maths: determinant | AMPS_Hard, in livebench/math | 1 | 1 | 24 June 2024 |
| Maths: greatest common divisor | AMPS_Hard, in livebench/math | 1 | 1 | 31 August 2024 |
| Maths: sample variance | AMPS_Hard, in livebench/math | 0 | 1 | 31 August 2024 |
| Maths: sample standard deviation | AMPS_Hard, in livebench/math | 0 | 1 | 31 August 2024 |
| Table matching | tablejoin, in livebench/data_analysis | 2 | 3 | 24 June 2024 |
LiveBench builds all 3 types itself. It generates the logic puzzles and the maths questions by program. It builds the table-matching questions from recent datasets on Kaggle and Socrata.
HIBN does not use LiveBench questions that come from maths competitions, coding sites, news articles or film synopses.
What HIBN changed
HIBN did not change any number, clue, table or correct answer. It made these edits to the wording:
- Logic puzzles: HIBN removed the request to reason step by step. The answer format is unchanged.
- Maths: in the determinant questions, HIBN turned the characters
\ninto real line breaks. In the determinant and greatest common divisor questions, it wrote\boxedwith one backslash where the source had two. The two statistics questions are unchanged. - Table matching: HIBN turned the characters
\ninto real line breaks.
HIBN also adds its own instructions at the top of each check and a label above each question. Methodology shows the instructions.
How a HIBN check differs from the way LiveBench is run
A HIBN score cannot be compared with any number on the LiveBench leaderboard. The two are produced in different ways:
In LiveBench’s files, each question stands alone, as its own item. LiveBench’s paper says: “For all models and tasks, we perform single-turn evaluation with temperature 0, unless otherwise noted in the model card.”
In a HIBN check, a person sends 6 or 10 questions in one message, through whatever app they use. The message asks for final answers only. HIBN cannot set or see the temperature, a setting that controls how much an AI’s replies vary. The app sets it, or the person can when they use an API that allows it.
LiveBench’s datasheet counts 960 questions. A HIBN check has 6 or 10, all published in 2024.
How HIBN’s scoring differs
HIBN’s scoring code is its own version of LiveBench’s scoring code for these 3 question types. It is based on LiveBench’s code at commit 8f8e5c381a16.
For maths answers, whenever its own comparison does not confirm a match, LiveBench’s code asks an AI model to decide. HIBN’s version has no such step. When HIBN’s code cannot decide, it records “Couldn’t be scored” and the check gets no overall score.
LiveBench’s code compares maths answers with a symbolic maths library. HIBN’s compares them as text, then as numbers. It keeps LiveBench’s own tolerance: two numbers count as equal when they are less than 0.001 apart.
In a test run recorded on 1 October 2026, HIBN’s version and LiveBench’s own code scored the same 270 test answers: for each of the 16 questions, the correct answer written in several ways, partly right and wrong answers, and empty or unreadable replies. They gave the same question score on 247. On the other 23, all maths answers, LiveBench’s code could not decide without its AI step, which was off for the test, and HIBN’s version scored none of them as correct. No one outside HIBN has reviewed HIBN’s version, and it is not published.
The score for a whole check is HIBN’s own rule: the plain average of 6 or 10 question scores. It is not a LiveBench score.
How old the questions are
LiveBench published these questions in 2024, and they have been public since. An AI model trained after that may have seen them, or their answers.
LiveBench has added newer questions since then, and it holds its newest ones back from public release. HIBN’s question sets are fixed, so the chance that an AI has seen them rises with time.
Licence and attribution
LiveBench’s datasheet says the benchmark is “distributed under the Apache License 2.0”. It also says: “There are no copyrights on the data.”
HIBN uses the questions on that basis. Its edits are listed above. The licence text is at Apache License 2.0.
HIBN’s scoring code is adapted from the scoring code in LiveBench’s repository. That repository’s licence file carries the text of the Apache License 2.0.
“LiveBench” is the name of that project. HIBN uses the name only to say where its questions and its scoring code come from.
To cite LiveBench, cite the paper named above. If you quote a HIBN number, cite HIBN and say that the questions come from LiveBench.
If you work on LiveBench and want something on this page corrected, write to support@haveibeennerfed.ai.