Some_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 5 days agoGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comexternal-linkmessage-square208linkfedilinkarrow-up1794arrow-down119cross-posted to: fuck_ai@lemmy.worldcollapse@lemmy.zipusa@lemmy.mlhackernews@lemmy.bestiver.seusa@midwest.socialtechnology@lemmit.online
arrow-up1775arrow-down1external-linkGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comSome_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 5 days agomessage-square208linkfedilinkcross-posted to: fuck_ai@lemmy.worldcollapse@lemmy.zipusa@lemmy.mlhackernews@lemmy.bestiver.seusa@midwest.socialtechnology@lemmit.online
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up4·5 days agoIs that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
minus-squareAsafum@lemmy.worldlinkfedilinkEnglisharrow-up3·5 days agoIt’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
minus-squareDamage@feddit.itlinkfedilinkEnglisharrow-up3·5 days agoMy framework 13 with shared RAM runs qwen quite well
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well