• Coriza@lemmy.world
    link
    fedilink
    English
    arrow-up
    19
    arrow-down
    3
    ·
    2 days ago

    The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.

    So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

    • KingRandomGuy@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      20 hours ago

      Keep in mind many of the optimizations you’re talking about (LRU caching of experts, for example) are only really relevant at the single user local inference scale. As in, an individual wants to run a big model on their machine, but they don’t have enough VRAM to fit the model and KV cache. Accordingly, you’re basically only describing hobbyist and research projects, which aren’t really representative of the AI inference industry as a whole.

      Commercial inference keeps everything resident in VRAM, so expert caching isn’t necessary. So these things won’t help costs. A lot of other low hanging fruit (like hierarchical KV cache) has also existed for a long time for production-ready inference engines.

      • Coriza@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        20 hours ago

        I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.

        • KingRandomGuy@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          4 hours ago

          It still won’t help for a couple of reasons. For one, VRAM is fairly abundant on commercial deployments. Even for big models, a company is probably deploying on 1-2 nodes of 8x H200 or newer (hence several TB of VRAM).

          But more importantly, commercial inference relies on heavy concurrency. So even if some experts are uncommon, with a lot of concurrent users, they will still fire frequently enough for the performance difference to be felt. And in my own experience, expert use isn’t uniform but it isn’t particularly biased either. This is especially tough since high-concurrency inference can actually be fairly compute bound, but this expert caching system either starves the system of bandwidth (if you require compute on the GPU, then you’re stuck with PCIe speeds which are tiny compared to HBM and even DRAM), or you’re starved of compute (if you do compute on the CPU).

          It’s nice for local inference, but yeah, not representative of commercial inference.

    • wewbull@feddit.uk
      link
      fedilink
      English
      arrow-up
      11
      ·
      2 days ago

      That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

      It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.

      • Coriza@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        21 hours ago

        Yes, I believe it was for a Mixture of experts model, which just goes to show how naive IA implementations are at this point, The cache was not solving any hard problem and yet for some reason it was not only not already standard practice but also somehow a notable achievement.