Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.
The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.
So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.
Keep in mind many of the optimizations you’re talking about (LRU caching of experts, for example) are only really relevant at the single user local inference scale. As in, an individual wants to run a big model on their machine, but they don’t have enough VRAM to fit the model and KV cache. Accordingly, you’re basically only describing hobbyist and research projects, which aren’t really representative of the AI inference industry as a whole.
Commercial inference keeps everything resident in VRAM, so expert caching isn’t necessary. So these things won’t help costs. A lot of other low hanging fruit (like hierarchical KV cache) has also existed for a long time for production-ready inference engines.
I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.
It still won’t help for a couple of reasons. For one, VRAM is fairly abundant on commercial deployments. Even for big models, a company is probably deploying on 1-2 nodes of 8x H200 or newer (hence several TB of VRAM).
But more importantly, commercial inference relies on heavy concurrency. So even if some experts are uncommon, with a lot of concurrent users, they will still fire frequently enough for the performance difference to be felt. And in my own experience, expert use isn’t uniform but it isn’t particularly biased either. This is especially tough since high-concurrency inference can actually be fairly compute bound, but this expert caching system either starves the system of bandwidth (if you require compute on the GPU, then you’re stuck with PCIe speeds which are tiny compared to HBM and even DRAM), or you’re starved of compute (if you do compute on the CPU).
It’s nice for local inference, but yeah, not representative of commercial inference.
That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.
It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.
Yes, I believe it was for a Mixture of experts model, which just goes to show how naive IA implementations are at this point, The cache was not solving any hard problem and yet for some reason it was not only not already standard practice but also somehow a notable achievement.
The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.
So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.
Keep in mind many of the optimizations you’re talking about (LRU caching of experts, for example) are only really relevant at the single user local inference scale. As in, an individual wants to run a big model on their machine, but they don’t have enough VRAM to fit the model and KV cache. Accordingly, you’re basically only describing hobbyist and research projects, which aren’t really representative of the AI inference industry as a whole.
Commercial inference keeps everything resident in VRAM, so expert caching isn’t necessary. So these things won’t help costs. A lot of other low hanging fruit (like hierarchical KV cache) has also existed for a long time for production-ready inference engines.
I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.
It still won’t help for a couple of reasons. For one, VRAM is fairly abundant on commercial deployments. Even for big models, a company is probably deploying on 1-2 nodes of 8x H200 or newer (hence several TB of VRAM).
But more importantly, commercial inference relies on heavy concurrency. So even if some experts are uncommon, with a lot of concurrent users, they will still fire frequently enough for the performance difference to be felt. And in my own experience, expert use isn’t uniform but it isn’t particularly biased either. This is especially tough since high-concurrency inference can actually be fairly compute bound, but this expert caching system either starves the system of bandwidth (if you require compute on the GPU, then you’re stuck with PCIe speeds which are tiny compared to HBM and even DRAM), or you’re starved of compute (if you do compute on the CPU).
It’s nice for local inference, but yeah, not representative of commercial inference.
That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.
It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.
Yes, I believe it was for a Mixture of experts model, which just goes to show how naive IA implementations are at this point, The cache was not solving any hard problem and yet for some reason it was not only not already standard practice but also somehow a notable achievement.