• wewbull@feddit.uk
    link
    fedilink
    English
    arrow-up
    11
    ·
    2 days ago

    That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

    It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.

    • Coriza@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      20 hours ago

      Yes, I believe it was for a Mixture of experts model, which just goes to show how naive IA implementations are at this point, The cache was not solving any hard problem and yet for some reason it was not only not already standard practice but also somehow a notable achievement.