Some_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 5 days agoGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comexternal-linkmessage-square201fedilinkarrow-up1687arrow-down118cross-posted to: [email protected]
arrow-up1669arrow-down1external-linkGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comSome_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 5 days agomessage-square201fedilinkcross-posted to: [email protected]
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up4·5 days agoIs that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
minus-squareAsafum@lemmy.worldlinkfedilinkEnglisharrow-up3·4 days agoIt’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
minus-squareDamage@feddit.itlinkfedilinkEnglisharrow-up3·4 days agoMy framework 13 with shared RAM runs qwen quite well
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well