eicker@lemmy.world to Technology@lemmy.worldEnglish · 1 day agoOpenAI Hoarding Tens Of Thousands Of Apple Mac mini And Mac Studio Devices, As ASUS And MSI Burn Through Their Entire First Batch Of NVIDIA RTX Spark Chip And Beg For More.wccftech.comexternal-linkmessage-square78linkfedilinkarrow-up1312arrow-down12
arrow-up1310arrow-down1external-linkOpenAI Hoarding Tens Of Thousands Of Apple Mac mini And Mac Studio Devices, As ASUS And MSI Burn Through Their Entire First Batch Of NVIDIA RTX Spark Chip And Beg For More.wccftech.comeicker@lemmy.world to Technology@lemmy.worldEnglish · 1 day agomessage-square78linkfedilink
minus-squareLydia_K@lemmy.worldlinkfedilinkEnglisharrow-up7·19 hours agohttps://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
minus-squareArchAengelus@lemmy.dbzer0.comlinkfedilinkEnglisharrow-up2·18 hours agoUpgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
minus-squareLydia_K@lemmy.worldlinkfedilinkEnglisharrow-up2·17 hours agoI plan to once there is a version with turboquant and MTP as that huge context window is key.
minus-squareboonhet@sopuli.xyzlinkfedilinkEnglisharrow-up1·19 hours agoA used 3090 is like 2-3k though, IF you can find one :| A month of claude is like 20 EUR. A month of opencode go is half that, but you get less usage.
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.
Upgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.
I plan to once there is a version with turboquant and MTP as that huge context window is key.
A used 3090 is like 2-3k though, IF you can find one :| A month of claude is like 20 EUR. A month of opencode go is half that, but you get less usage.