About

About LiveBench

The questions in Quick and Deep come from LiveBench, a public benchmark for AI language models. HIBN is not affiliated with LiveBench and is not endorsed by it. A HIBN score is not a LiveBench score.

What LiveBench is

LiveBench describes itself as “a benchmark for LLMs designed with test set contamination and objective evaluation in mind”. LLMs are the large language models behind AI chat apps.

It was created by Colin White and 17 co-authors. They are from Abacus.AI, New York University, Nvidia, the University of Maryland, the University of Southern California and Columbia University. LiveBench’s datasheet says: “This benchmark was funded by Abacus.AI.”

The paper is “LiveBench: A Challenging, Contamination-Limited LLM Benchmark”. LiveBench’s README says: “LiveBench appeared as a Spotlight Paper in ICLR 2025.”

LiveBench publishes a leaderboard of AI models. It adds new questions regularly, because questions that stay public can end up in the data that AI models are trained on.

LiveBench website · LiveBench on GitHub · The LiveBench paper

HIBN read these LiveBench sources on 2 October 2026: the paper (version 2), the README, the datasheet, the licence file and the maths scoring code.

What HIBN uses

HIBN uses 16 LiveBench questions of 3 types. It took them from LiveBench’s public datasets on Hugging Face.

HIBN’s nameLiveBench’s name and datasetIn QuickIn DeepLiveBench release
Logic puzzlezebra_puzzle, in livebench/reasoning2325 November 2024
Maths: determinantAMPS_Hard, in livebench/math1124 June 2024
Maths: greatest common divisorAMPS_Hard, in livebench/math1131 August 2024
Maths: sample varianceAMPS_Hard, in livebench/math0131 August 2024
Maths: sample standard deviationAMPS_Hard, in livebench/math0131 August 2024
Table matchingtablejoin, in livebench/data_analysis2324 June 2024

LiveBench builds all 3 types itself. It generates the logic puzzles and the maths questions by program. It builds the table-matching questions from recent datasets on Kaggle and Socrata.

HIBN does not use LiveBench questions that come from maths competitions, coding sites, news articles or film synopses.

What HIBN changed

HIBN did not change any number, clue, table or correct answer. It made these edits to the wording:

  • Logic puzzles: HIBN removed the request to reason step by step. The answer format is unchanged.
  • Maths: in the determinant questions, HIBN turned the characters \n into real line breaks. In the determinant and greatest common divisor questions, it wrote \boxed with one backslash where the source had two. The two statistics questions are unchanged.
  • Table matching: HIBN turned the characters \n into real line breaks.

HIBN also adds its own instructions at the top of each check and a label above each question. Methodology shows the instructions.

How a HIBN check differs from the way LiveBench is run

A HIBN score cannot be compared with any number on the LiveBench leaderboard. The two are produced in different ways:

In LiveBench’s files, each question stands alone, as its own item. LiveBench’s paper says: “For all models and tasks, we perform single-turn evaluation with temperature 0, unless otherwise noted in the model card.”

In a HIBN check, a person sends 6 or 10 questions in one message, through whatever app they use. The message asks for final answers only. HIBN cannot set or see the temperature, a setting that controls how much an AI’s replies vary. The app sets it, or the person can when they use an API that allows it.

LiveBench’s datasheet counts 960 questions. A HIBN check has 6 or 10, all published in 2024.

How HIBN’s scoring differs

HIBN’s scoring code is its own version of LiveBench’s scoring code for these 3 question types. It is based on LiveBench’s code at commit 8f8e5c381a16.

For maths answers, whenever its own comparison does not confirm a match, LiveBench’s code asks an AI model to decide. HIBN’s version has no such step. When HIBN’s code cannot decide, it records “Couldn’t be scored” and the check gets no overall score.

LiveBench’s code compares maths answers with a symbolic maths library. HIBN’s compares them as text, then as numbers. It keeps LiveBench’s own tolerance: two numbers count as equal when they are less than 0.001 apart.

In a test run recorded on 1 October 2026, HIBN’s version and LiveBench’s own code scored the same 270 test answers: for each of the 16 questions, the correct answer written in several ways, partly right and wrong answers, and empty or unreadable replies. They gave the same question score on 247. On the other 23, all maths answers, LiveBench’s code could not decide without its AI step, which was off for the test, and HIBN’s version scored none of them as correct. No one outside HIBN has reviewed HIBN’s version, and it is not published.

The score for a whole check is HIBN’s own rule: the plain average of 6 or 10 question scores. It is not a LiveBench score.

How old the questions are

LiveBench published these questions in 2024, and they have been public since. An AI model trained after that may have seen them, or their answers.

LiveBench has added newer questions since then, and it holds its newest ones back from public release. HIBN’s question sets are fixed, so the chance that an AI has seen them rises with time.

Licence and attribution

LiveBench’s datasheet says the benchmark is “distributed under the Apache License 2.0”. It also says: “There are no copyrights on the data.”

HIBN uses the questions on that basis. Its edits are listed above. The licence text is at Apache License 2.0.

HIBN’s scoring code is adapted from the scoring code in LiveBench’s repository. That repository’s licence file carries the text of the Apache License 2.0.

“LiveBench” is the name of that project. HIBN uses the name only to say where its questions and its scoring code come from.

To cite LiveBench, cite the paper named above. If you quote a HIBN number, cite HIBN and say that the questions come from LiveBench.

If you work on LiveBench and want something on this page corrected, write to support@haveibeennerfed.ai.