• CompactFlax@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    46
    ·
    edit-2
    2 days ago

    “Look how low our cost of inference is”

    “Pay no attention to the marketing budget that exceeds Coca-Cola’s for a small fraction of their revenues”

    What technical, fundamental reason is there for the crash in price? The article just accepts the MSRP as fact. It’s established fact that retail prices can be dropped below cost in order to establish market dominance. The cost of training can indeed be spread over time but it’s not spread across enough time (between model releases)The inference cost doesn’t actually drop in reality.

    • humanspiral@lemmy.ca
      link
      fedilink
      English
      arrow-up
      1
      ·
      7 hours ago

      What technical, fundamental reason is there for the crash in price?

      Smaller models are inherently faster, and use less gpus to fit inside, and less gpu time per training step. Newer models are smarter at smaller size than older models, and so also come up with correct answer in fewer tokens.

      There are 15 labs accross the world competing without patent restriction for software building. It’s orders of magnitude faster than Moore’s law, because its highly competitive, and software has massive “compiler” resources thrown at it, in investor/government frenzy.

    • Grimy@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      12
      ·
      edit-2
      2 days ago

      There are constantly new techniques being developed to speed up inference or reduce model size post training. There’s about a hundred different levers to pull that play on inference cost, and some of them don’t have much of an impact on quality.

      In the end, it’s probably simply because of competition. I don’t understand why everybody assumes they are running these at cost API wise.

      Edit: here’s a chart from ars technica. The cost of revenue is clearly lower than the actual revenue. They aren’t running inference at a lost or it would be higher. They aren’t profitable because they are spending all that money on capturing the market as quick as they can. For fucks sake, that includes all the free accounts as well running inference. How can anyone think a paying customer is getting it at less than cost?

      And yes, there have been advancement made. The cost of inference isn’t some static number that can never go down. Stop believing them when they tell you there’s no profit to be made. It’s the same playbook as a dozen other companies, they literally get rewarded for it when it comes tax time.

      • Jiral@lemmy.world
        link
        fedilink
        English
        arrow-up
        21
        ·
        2 days ago

        Yet all major players are still incapable of covering their costs and have to resort to all sorts of nasty accounting tricks to make it apear otherwise. How so?

        • Grimy@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          10
          ·
          2 days ago

          Because of the massive investments into datacenters. Why is everyone so quick to drink the kool-aid. They are spending stupid amounts of money to capture the market and force everyone into a subscription service, not because it cost money to actually run the models. They are all switching their models to fine tuned quants after the first two weeks as well.

          The open source community has found dozens of ways to lower costs, do you really think these big companies aren’t using the same tricks and haven’t developed even more advanced techniques?

          • Jiral@lemmy.world
            link
            fedilink
            English
            arrow-up
            11
            ·
            2 days ago

            It doesn’t matter what tricks they supposedly use afterwards, the data centets that are built need to run successfully or the bubble pops and then those companies will fall like lead. In other words for those data centers to be successfull the entire real economy has to spend crazy amounts of money on stuff that needs crazy amounts of compute.

            Open AI and Anthropic would really need some good business numbers, right now. The claim that they are just deliberately making their business case much worse than it is, is insane.

            • Grimy@lemmy.world
              link
              fedilink
              English
              arrow-up
              7
              arrow-down
              2
              ·
              edit-2
              2 days ago

              The inference cost doesn’t actually drop in reality.

              My comment was about this which is just completely wrong.

              OpenAI is the scummiest company on earth but everyone is ready to believe them when they say shit like "our 20$ plan lets you run up the equivalent of a billion dollars in API fees. Giggles. It’s a good plan for you but not for us ;) "

              The main thing bringing up costs is stuff like Sam Altman spending all the “inference” money on raw wafers just to strangle consumer GPU sales.

              It’s not a good business model because of how cheap inference actually is. Once you have open models and the moat is broken, the only thing left to do is regulatory capture, upping the demand for the hardware and little charades like these.