About

Methodology

Method version 1 · last updated 2 October 2026

In a check, you give a fixed set of questions to your AI by copy and paste. HIBN scores the reply with fixed code. If it can find and score every answer, it gives a score out of 100.

A score describes how an AI did on a small set of questions, with one setup, at the time of the check. It does not measure an AI’s overall quality. It cannot show that an AI was changed, or why a score moved.

What a check is

You copy the questions from HIBN and paste them into your AI as one message. You paste the AI’s first reply back. HIBN finds each answer in the reply and scores it.

HIBN does not connect to your AI. It sees only the reply you paste, and nothing else from your chat. It cannot see your account, your settings or which model answered.

A check applies to whatever you paste the questions into: a chat app, a desktop or mobile app, a coding tool or an API. You tell HIBN which, and HIBN records what you tell it.

Each check takes one reply. A second attempt is a new check.

The questions

HIBN has 2 fixed question sets. Quick has 6 questions and Deep has 10, and they share none. Every Quick check uses the same 6 questions in the same order. Every Deep check uses the same 10.

Question typeIn QuickIn DeepWhat the AI has to do
Logic puzzle23Work out who has which attribute from a list of clues, then give several answers about the result.
Maths24Give one exact value. Quick asks for a determinant and a greatest common divisor. Deep asks for those two kinds, plus a sample variance and a sample standard deviation.
Table matching23Match the columns of one data table to the columns of another.

All 16 questions come from LiveBench, a public benchmark for AI language models. LiveBench published them in its releases of 24 June 2024, 31 August 2024 and 25 November 2024. About LiveBench says what HIBN took and what it changed.

HIBN kept 8 questions from 2 earlier versions of its question sets: every question in them that LiveBench had generated itself. No fixed rule is recorded for how those earlier versions chose their questions. HIBN picked the other 8 by fixed rules. Of these, 6 are the longest eligible question of their kind within a length limit. The limit is 1,550 to 4,000 characters, depending on the question type and the set. The other 2 are greatest common divisor questions, each picked because its numbers include the largest among the eligible questions. The 16 are not a random sample of LiveBench, and they do not represent it.

Deep is not known to be harder than Quick. Its questions are longer: about 2,400 characters each on average, against about 1,200 for Quick.

What the AI is asked to do

You send the questions as one message. A Quick check begins with the instructions below. A Deep check uses the same words, with its own name, 10 questions and the labels D01 to D10.

haveibeennerfed.ai — Quick check (HIBN-LBG-1-QUICK, form A)

Answer all 6 independent questions using only the information supplied. Do not browse, run code, use a calculator or use other tools.
Reply with sections headed Q01 through Q06, in order, each in that question's requested final-answer format. Give final answers only, without step-by-step working. Instructions about "the end of your answer" refer to that question's section. Do not wrap the whole reply in JSON.
If you can't answer a question, say so in its section rather than leaving it out.

HIBN cannot tell whether the AI followed these instructions. If web search, tools that run code, or memory were on, the AI may have used them.

The questions ask for final answers only, and HIBN removed the request to reason step by step from the logic puzzles. An AI with a thinking setting can still reason before it replies.

Putting all the questions in one message may itself affect how an AI does. HIBN has not measured that effect.

How HIBN finds the answers

HIBN looks for each label at the start of a line. The labels are Q01 to Q06 for Quick, D01 to D10 for Deep, and T01 onwards for a Custom Pack. A label can be plain, bold, a Markdown heading, or set between equals signs. The answer can follow a colon on the same line, or start on the next line.

The text between one label and the next is that question’s answer. Text before the first label is ignored.

If a label is not in the reply, has nothing after it, or appears more than once, HIBN records that answer as missing.

Before you get your result, the page shows which labels it found. HIBN stores the reply exactly as you pasted it, with a SHA-256 fingerprint of the text.

How each question is scored

Fixed code on HIBN’s server scores each answer. No AI model takes part in scoring. Each answer that can be scored gets a question score from 0 to 1.

The code is HIBN’s own version of LiveBench’s scoring code for these question types. HIBN’s version is not published. About LiveBench lists the differences.

Logic puzzles

Each puzzle asks for several answers in a fixed order, inside <solution> tags. HIBN compares them in order. It ignores capital letters, treats a hyphen as a space, and counts an answer as right if it contains the correct answer. HIBN is sure to read the answers only when the tags and the answers are on one line.

Half of the question score comes from the share of answers that are right. The other half is given only when every answer is right. For example, when a puzzle asks for 4 answers and 3 are right, the question scores 0.375. With all 4 right, it scores 1.

Maths

Each question asks for one final answer in \boxed{} form. The answer is right if it is the same text as the correct answer. It is also right if both work out to numbers less than 0.001 apart. Right scores 1 and wrong scores 0.

If HIBN finds a final answer with a number in it but can’t work out its value, the question is recorded as “Couldn’t be scored”. It is not counted as incorrect. A final answer with no number in it scores 0.

Table matching

The answer is a list of column pairs. HIBN compares it with the correct pairs and gives an F1 score from 0 to 1. F1 is 1 when every pair is right, and it falls with each wrong or missing pair. The scoring code rounds F1 to 2 decimal places, and that rounded value is the question score.

Answers HIBN can’t read

If HIBN cannot find a final answer in a format it can read, the question scores 0. HIBN usually records “Answer format not recognised”. For a logic puzzle it records “Incorrect” when it finds the <solution> tags but cannot read the answers between them. That happens, for example, when the closing tag is on a line of its own. A question can score 0 in these ways even when the answer itself was right.

The outcomes

OutcomeQuestion scoreWhen
Correct1The answer matches.
Partly correctBetween 0 and 1Part of a logic puzzle or table-matching answer is right.
Incorrect0HIBN found an answer and it does not match. For a logic puzzle, also when HIBN found the answer tags but could not read the answers.
Answer format not recognised0HIBN could not find a final answer in a format it can read.
Refused (you told us)0The answer is missing and you told HIBN the AI refused.
Not answered (you told us)0The answer is missing and you told HIBN the AI didn’t answer.
MissingNo question scoreHIBN could not find the answer, and you left it as missing.
Couldn’t be scoredNo question scoreHIBN found an answer, and its scoring code could not tell whether it is correct. For maths, this happens when the code can’t work out the value of a number.

How the score is worked out

The score is the average of the question scores, multiplied by 100 and rounded to a whole number. A half rounds up. Every question has the same weight. Before the average, the only rounding is of the table-matching F1, to 2 decimal places. The result table shows question scores to 2 decimal places.

An example for a Quick check: question scores of 1, 0.375, 1, 0, 0.67 and 1 average 0.674. The score is 67.

Beside the score, HIBN shows how many questions were fully correct. A question is fully correct when its question score is 1.

When a check has no score

If any answer is missing, or any answer couldn’t be scored, the check has no overall score. The other questions still show their outcomes.

HIBN does not treat a missing answer as incorrect. An answer can be missing because the reply was cut off, pasted in part, or written without labels.

A missing answer scores 0 only when you tell HIBN that the AI refused or didn’t answer. HIBN marks those answers “you told us”.

HIBN never decides from the text that an AI refused. If the AI writes under a label that it can’t answer, HIBN scores that text like any other answer, and it scores 0.

A check with no score is not counted in Analytics and cannot be compared with another check. So a median describes only the checks that produced a score. HIBN does not publish how many checks had no score.

What one question is worth

The question sets are small, so one answer moves the score by many points. One Quick question is worth up to 17 points. One Deep question is worth up to 10.

An AI can answer the same question differently on two attempts. HIBN has not measured how much scores vary from one check to the next. So it cannot say how big a change must be before chance is unlikely.

Treat any one change as possibly chance. More comparable checks show whether a difference repeats.

What you tell HIBN, and what it can’t confirm

The setup is what you choose: provider, app, model and level. HIBN cannot confirm it.

The lists of apps, models and levels come from each provider’s own documentation. One level, for 2 models in Codex and Work, comes from the Codex software itself. Where a provider says a model has levels but doesn’t list them, the level is Unknown. HIBN last looked at them on 3 October 2026. If your app shows something the lists don’t have, choose Other and type it.

Refusals and non-answers are recorded only when you say so.

You start and stop the stopwatch. The time includes copying and switching apps. It is not the AI’s speed, and it does not affect the score.

Location is estimated from your connection when the check starts. On the home page or a pack’s page you can change it or choose not to say. HIBN then stores only your choice; otherwise it stores the estimate.

Comparing two checks

HIBN shows a change in points only when two checks are comparable. All of these must be true:

  • They used the same question set. For a Custom Pack, they used the same version of the pack.
  • They were scored by the same version of the scoring code, and their scores were worked out by the same version of the score calculation.
  • They have the same setup, with nothing set to Unknown or Other.
  • Both have a score.
  • They were started at different times.

Unknown never matches Unknown. Two checks that both say Unknown for the model may have used different models.

A setting your app doesn’t have is not treated as Unknown. Microsoft Copilot has no model picker, so two checks from the same Copilot app in the same mode are comparable.

The change is the later score minus the earlier score. A result page compares the check with the most recent comparable one among your last 20 earlier finished checks of the same questions. In My Results you can pick any two results, and HIBN shows a change only if they are comparable.

A change does not show that the AI was changed, and it does not show why the score moved. The AI may have answered differently by chance. Something may have differed on your side or on the provider’s side. HIBN cannot tell these apart.

Grok’s web chat and mobile app don’t let you choose the model, so HIBN records none. xAI’s pages name modes, which HIBN records: Auto, Fast, Multi-agent and Build. Two checks compare only when they used the same mode. Checks made before 4 October 2026 have no mode recorded. HIBN still compares two of those from the same app, and beside the change the result page says that Grok’s mode isn’t recorded.

Custom Packs

A Custom Pack is a set of 1 to 20 questions written by a person, with the answer they expect for each. HIBN does not confirm that the expected answers are right.

The author picks one rule for each question: exact text, text in any capitalisation, or a number. An answer that matches under the rule scores 1. An answer that doesn’t match scores 0. As with Quick and Deep, a missing answer means no overall score, unless the person says the AI refused or didn’t answer.

A pack tells the AI to reply with only the answer. Under the text rules, an answer with more words or lines than the expected answer scores 0: the pack’s author sees it as “Answer format not recognised”, and anyone else sees “Incorrect”. Under the number rule, anything that is not one plain number is recorded as “Answer format not recognised”.

Pack checks are never counted in Analytics. The Pack guide covers how to write questions that can be scored.

How Analytics is produced

Analytics counts Quick and Deep checks that have a score. It leaves out a check when its owner has deleted it or HIBN has left it out. It counts checks, not people: one person can run many checks, and each one counts. A total for a past period can fall when a check is deleted or left out.

HIBN can leave a check out of Analytics, for example one that looks automated or abusive. A person at HIBN decides, and HIBN records who decided, when and why. The person who ran the check keeps their result. HIBN does not publish how many checks it has left out.

Quick and Deep are always shown separately. Their questions differ, so their scores are never combined.

Each total is an exact count. Counts by country are shown as ranges: fewer than 20, 20 to 49, 50 to 99, 100 to 499, and 500 or more.

A setup is listed only when it has at least 20 counted checks in the period. Its count is shown as a range: 20 to 49, 50 to 99, 100 to 499, or 500 or more. Its median is shown only when nothing in the setup is Unknown. Checks with Other anywhere in the setup are not grouped into setups.

The median is the middle score when a setup’s scores are put in order. With an even number of checks, it is the average of the two middle scores, rounded to a whole number. A half rounds up.

“Latest 4 weeks” means the current week so far and the 3 weeks before it. Weeks start on Monday at 00:00 UTC.

People choose to run checks, and they report their own setups. The numbers do not represent all users of any AI, and Analytics is not a ranking.

HIBN does not publish comparisons of scores between countries.

Limits

  1. Small question sets. 6 or 10 questions of 3 types can’t describe an AI in general. The Quick and Deep questions do not cover writing, coding, research, factual knowledge, tool use or safety.
  2. Public questions. The questions were published in 2024. An AI model trained after that may have seen them, or their answers. A high score may partly reflect that.
  3. The same questions every time. Each question set has one fixed version. An app that remembers earlier chats could carry answers from one check to the next.
  4. The setup is what you tell HIBN. HIBN records it and cannot confirm it.
  5. No control over tools. The questions ask the AI not to browse or run code. HIBN cannot tell whether it complied.
  6. Variation between checks. HIBN has not measured how much its scores vary from one check to the next.
  7. One message. All the questions go in one message and ask for final answers only. In LiveBench’s own files, each question stands alone. A HIBN score is not a LiveBench score.
  8. Answer format. If HIBN cannot find a final answer in a format it can read, the question scores 0. That is so even if the answer itself was right. If HIBN finds a maths answer with a number in it but cannot work out its value, the check has no overall score.
  9. A recognisable check. Every check begins with a line that names haveibeennerfed.ai, and the questions do not change. A provider’s system could recognise a HIBN check and treat it differently from other messages. HIBN has no way to tell whether that happens.
  10. Pasted replies. HIBN cannot confirm that a reply came from the AI named, or that it was pasted unchanged.
  11. People who chose to take part. Analytics counts the checks people chose to run, and one person can run many. It is not a survey.
  12. Not validated. HIBN has not tested whether its scores follow the quality people notice in everyday use.
  13. No independent review. Only HIBN has tested the question sets and the scoring code.

Versions

Every result records the versions it was scored with. The current versions are:

Question sets
HIBN-LBG-1-QUICK-A (Quick) · HIBN-LBG-1-DEEP-A (Deep)
Scoring code
hibn-lbg-graders-1
Score calculation
hibn-score-1
LiveBench scoring code it is based on
commit 8f8e5c381a16
Provider lists
provider-catalogue-2026-10-03b, as at 3 October 2026

From launch, HIBN gives its scoring code a new version whenever it changes how Quick and Deep answers are scored, and does not compare results across that change. The score calculation has its own version, and HIBN does not compare results across a change to it either.

Changes to this method

  • Version 1, 2 October 2026: first public version.

Who has reviewed this

HIBN has tested its question sets and scoring code itself. The stored correct answers score 100. A reply of “I don’t know” to every question scores 0. In a test run recorded on 1 October 2026, HIBN’s scoring code and LiveBench’s own scored the same 270 test answers: for each of the 16 questions, the correct answer written in several ways, partly right and wrong answers, and empty or unreadable replies. They gave the same question score on 247. On the other 23, all maths answers, LiveBench’s code could not decide without its AI step, which was off for the test, and HIBN’s code scored none of them as correct.

No one outside HIBN has reviewed HIBN’s choice of questions, its edits to them or its scoring code. The questions and their correct answers are LiveBench’s own.

To report an error in the method, use the Feedback page.