• boonhet@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    4
    ·
    1 day ago

    The tokens ARE language. They’re just words or parts of words.

    If you’re using a western LLM, English IS its strongest language. Otherwise it might be Chinese or English.

    There’s no “LLM native language” to convert to for efficiency. Maybe pseudocode or actual code for things where you need to disambiguate.

    • MangoCats@feddit.it
      link
      fedilink
      English
      arrow-up
      1
      ·
      1 day ago

      There’s no “LLM native language” to convert to for efficiency.

      Yet. With 1000x as much bot-content being generated on the web as human generated content, the “bot spoken” training set will grow rather quickly, and I see no reason for it not to evolve in a similar way to how human spoken languages evolve.

      • boonhet@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        1
        ·
        21 hours ago

        What’s the point in coming up with an entire new language for them if they can just write English and humans will also understand it?

        Kinda pointless to have LLMs optimised to produce garbage humans can’t understand because at the end, your goal is still to produce something humans will consume.