Evaluating 35 open-weight models across three context lengths (32K, 128K, 200K), four temperatures, and three hardware platforms—consuming 172 billion tokens across more than 4,000 runs—we find that the answer is “substantially, and unavoidably.” Even under optimal conditions—best model, best temperature, temperature chosen specifically to minimize fabrication—the floor is non-zero and rises steeply with context length. At 32K, the best model (GLM 4.5) fabricates 1.19% of answers, top-tier models fabricate 5–7%, and the median model fabricates roughly 25%.

  • jacksilver@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    4 months ago

    Just for context, this is the error rate when the right answer is provided to the LLM in a document. This means that even when the answer is being handed to the LLM they fail at the rates provided in the article/paper.

    Most people interacting with LLMs aren’t asking questions against documents, or the answer can not be directly inferred from the documents (asking the LLM to think about the materials in the documents).

    That means in most situations the error rate for the average user will be significantly higher.