Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.
There are constantly new techniques being developed to speed up inference or reduce model size post training. There’s about a hundred different levers to pull that play on inference cost, and some of them don’t have much of an impact on quality.
In the end, it’s probably simply because of competition. I don’t understand why everybody assumes they are running these at cost API wise.
Edit: here’s a chart from ars technica. The cost of revenue is clearly lower than the actual revenue. They aren’t running inference at a lost or it would be higher. They aren’t profitable because they are spending all that money on capturing the market as quick as they can. For fucks sake, that includes all the free accounts as well running inference. How can anyone think a paying customer is getting it at less than cost?
And yes, there have been advancement made. The cost of inference isn’t some static number that can never go down. Stop believing them when they tell you there’s no profit to be made. It’s the same playbook as a dozen other companies, they literally get rewarded for it when it comes tax time.
Yet all major players are still incapable of covering their costs and have to resort to all sorts of nasty accounting tricks to make it apear otherwise. How so?
Because of the massive investments into datacenters. Why is everyone so quick to drink the kool-aid. They are spending stupid amounts of money to capture the market and force everyone into a subscription service, not because it cost money to actually run the models. They are all switching their models to fine tuned quants after the first two weeks as well.
The open source community has found dozens of ways to lower costs, do you really think these big companies aren’t using the same tricks and haven’t developed even more advanced techniques?
It doesn’t matter what tricks they supposedly use afterwards, the data centets that are built need to run successfully or the bubble pops and then those companies will fall like lead. In other words for those data centers to be successfull the entire real economy has to spend crazy amounts of money on stuff that needs crazy amounts of compute.
Open AI and Anthropic would really need some good business numbers, right now. The claim that they are just deliberately making their business case much worse than it is, is insane.
The inference cost doesn’t actually drop in reality.
My comment was about this which is just completely wrong.
OpenAI is the scummiest company on earth but everyone is ready to believe them when they say shit like "our 20$ plan lets you run up the equivalent of a billion dollars in API fees. Giggles. It’s a good plan for you but not for us ;) "
The main thing bringing up costs is stuff like Sam Altman spending all the “inference” money on raw wafers just to strangle consumer GPU sales.
It’s not a good business model because of how cheap inference actually is. Once you have open models and the moat is broken, the only thing left to do is regulatory capture, upping the demand for the hardware and little charades like these.
There are constantly new techniques being developed to speed up inference or reduce model size post training. There’s about a hundred different levers to pull that play on inference cost, and some of them don’t have much of an impact on quality.
In the end, it’s probably simply because of competition. I don’t understand why everybody assumes they are running these at cost API wise.
Edit: here’s a chart from ars technica. The cost of revenue is clearly lower than the actual revenue. They aren’t running inference at a lost or it would be higher. They aren’t profitable because they are spending all that money on capturing the market as quick as they can. For fucks sake, that includes all the free accounts as well running inference. How can anyone think a paying customer is getting it at less than cost?
And yes, there have been advancement made. The cost of inference isn’t some static number that can never go down. Stop believing them when they tell you there’s no profit to be made. It’s the same playbook as a dozen other companies, they literally get rewarded for it when it comes tax time.
Yet all major players are still incapable of covering their costs and have to resort to all sorts of nasty accounting tricks to make it apear otherwise. How so?
Because of the massive investments into datacenters. Why is everyone so quick to drink the kool-aid. They are spending stupid amounts of money to capture the market and force everyone into a subscription service, not because it cost money to actually run the models. They are all switching their models to fine tuned quants after the first two weeks as well.
The open source community has found dozens of ways to lower costs, do you really think these big companies aren’t using the same tricks and haven’t developed even more advanced techniques?
It doesn’t matter what tricks they supposedly use afterwards, the data centets that are built need to run successfully or the bubble pops and then those companies will fall like lead. In other words for those data centers to be successfull the entire real economy has to spend crazy amounts of money on stuff that needs crazy amounts of compute.
Open AI and Anthropic would really need some good business numbers, right now. The claim that they are just deliberately making their business case much worse than it is, is insane.
My comment was about this which is just completely wrong.
OpenAI is the scummiest company on earth but everyone is ready to believe them when they say shit like "our 20$ plan lets you run up the equivalent of a billion dollars in API fees. Giggles. It’s a good plan for you but not for us ;) "
The main thing bringing up costs is stuff like Sam Altman spending all the “inference” money on raw wafers just to strangle consumer GPU sales.
It’s not a good business model because of how cheap inference actually is. Once you have open models and the moat is broken, the only thing left to do is regulatory capture, upping the demand for the hardware and little charades like these.