• Barbecue Cowboy@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    18
    ·
    4 days ago

    It’s kinda surprising,

    I know specifically where one of the big ones hosts its models and its not there, but I guess they could have infrastructure in there.

    • [object Object]@lemmy.ca
      link
      fedilink
      English
      arrow-up
      22
      ·
      4 days ago

      They’re oversold though, especially prompt caching and the parameter count war

      The US model is that they think more training compute and parameters will result in the winning model, while the Chinese are focusing on RL and parameters efficiency due to compute limits.

        • Dave.@aussie.zone
          link
          fedilink
          English
          arrow-up
          8
          ·
          4 days ago

          Always ready to try brute force first. And then some other, less palatable options if that doesn’t work.

      • percent@infosec.pub
        link
        fedilink
        English
        arrow-up
        5
        ·
        4 days ago

        The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.

        (It’s not super complex code, just some scripts that I would not have taken to time to write manually.)

        I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬

      • teslekova@lemmy.ml
        link
        fedilink
        English
        arrow-up
        2
        ·
        3 days ago

        Considering which country is better at building power stations, that’s a fascinating dichotomy.

        The US going for brute force when the brute force is more available in China… Priceless irony.

        • [object Object]@lemmy.ca
          link
          fedilink
          English
          arrow-up
          2
          ·
          3 days ago

          China doesn’t have the near the amount of compute resources, but they can run less efficient servers for cheaper, so it’s a wash.