• percent@infosec.pub
    link
    fedilink
    English
    arrow-up
    5
    ·
    5 days ago

    The efficiency of Chinese models really is impressive. I generated sooo much code yesterday with Qwen3.6 35B-A3B running on an RTX 5060 Ti 16GB (+ a little CPU offloading). It got the jobs done at ~50 tokens/sec.

    (It’s not super complex code, just some scripts that I would not have taken to time to write manually.)

    I’d love to upgrade to something with more VRAM, but even my current card has doubled in price since I bought it last year 😬