I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
Is it really feasible to run it 100% locally?
Yes. My Macbook Air (M2) released in 2022 can run many publicly available LLM models. The ouput is not as fast as using a large powerful datacenter, but for my local needs, I’m not in a hurry. I get about 17 to 30 tokens per second speed running 100% locally.
If yes to the previous: the software doesn’t come from nowhere
For Mac users, the interface comes from here . For the specific LLM models that is a separate question for each.
and ultimately still relies on gas-turbine-powered datacenters
DCs generating power on-site is a relatively new phenomon because existing grids are at capacity, so the only way to bring new DCs online is locally generating power at that DC, usually using gas turbines or even worse, diesel generators. Most if not all of the publicly available models for running on your own hardware were built before those gas-turnbine-generating DCs were a thing.
The public models people are running now have existed for a number of years are likely made on regular utility grid power which is whatever that nation and region uses.
and stolen IP and stolen personal data, no?
The Llama LLM is made by Meta, so probably yes for that one. Deepseek is from an AI research lab in China. QWEN is from Chinese company Alibaba. We don’t know for sure the inputs that created the Chinese models. US AI companies claim a number of the Chinese models are derived from American LLMs, but I haven’t seen (or looked for) proof of these claims.
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically
Likely they mean because you don’t have to pay a large LLM owner in a rent-seeking model to run LLMs, nor does a person’s use contribute to further development by those companies.
or true to FOSS philosophy, because … ???
With the open-weight models it doesn’t rely on a commercial license to use, and the interfaces can be truly open source.
I don’t know if 5 year olds are allowed to watch 1.5h videos, but if so: this one has all the pro and con arguments explained nicely and makes a point why it’s not a good idea to shame people for using LLMs.
You can run them locally, yes. There are models that can even run on phones, but usecase is limited. But it can only be considered ethical, if the training data used is listed or ethically sourced IMO.
AI bros on Lemmy will disagree with me, but most open weight models are still trained unethically i.e, theft. Most proponents of LLMs (who I talked to on bsky), who say local models are ethical, don’t fucking use it. They’re larping on socials about how awesome it is, but none of the ones I talked to are using it in their projects. They mess around, realise it is not as good as the “unethical” options, go right back to Claude
Open weight models Qwen, deepseek, mistral, and the Ollama stuff etc are unethical in normal people’s eyes, but “ethical” enough for AI bros.
From what I searched, there are very few that can be considered ethical - Olmo, Apertus, Starcoder(?). But idk anyone who uses these. My friend at IBM said they used Apertus, but it was nowhere near good as ChatGPT, so they no longer use Apertus now. And these models require minimum 6-8 GB VRAM for their lowest parameter model iirc.
Even the open-weight model bros are lobbying to redefine what ‘open-source AI’ means. That should give you a fair idea about people behind open-weight as well
I would unironically argue that a model primarily trained through distillation of closed frontier models, and then released open-weight with an open-source architecture, becomes “ethical” again.
Something something Robin Hood
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
I’m not holding this as a hard stance, just a current analysis of the technology and how it might affect my life in my circumstances. I think we’re well into a world where we need to be deciding if some technologies don’t suit our lives, because there certainly are even more harmful technologies to come.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
In leaving progress to the machines, in letting technology go forward on its own terms and selecting from it, with what seems to us excessive caution, modesty, or restraint, the limited though completely adequate implements of their cultures, is it possible that in thus opting not to move “forward” or not only “forward,” these people did in fact succeed in living in human history, with energy, liberty, and grace?
Always Coming Home - Stone Telling Part 3 by Ursula K. Le Guin
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
How is LLM output any more harmful than output from a human you don’t know? I would agree with you if one were to simply blindly accept anything an LLM gives you as factual. However, I am skeptical of what humans say too. Some are inaccurate out of carelessness, some out of malice. Critical thinking is the key to protection from both errand LLM output and fallible humans.
Additionally, this thread is about locally running LLMs. One of the most dangerous aspects of most large LLMs is the sycophancy where the LLM will try to tell you what you want to hear, even if it needs to creatively invent things that don’t exist or are not true. If you are not aware, running locally means you have all the controls on the model. You’re not subject to whatever settings a large hyperscaler sets up for you. This means, among other things, you can turn down the “temperature” setting, which is the lever that controls how creative an LLM is. Practically what this means is, if you set it to “0” you are allowing it no creative action. If you ask it a question it doesn’t have actual training on, it will tell you effectively “I don’t know” instead of making something up just to have an answer.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
The older I get the more disappointed I get in humanities large group decisions and actions.
LLMs regardless are awful. Running whatever model on your own hardware is fine, but it depends on if what you’re doing with it is creative or not. Writing, art, etc. are out of the question. Help with coding, pattern recognition on a large subset of data, etc. is fine and isn’t too far from stuff we used to use for those purposes anyways.
The issue at the end of the day isn’t AI. It’s where and how it’s used and, on a larger scale, what markets and environments are thrown into disarray because of idiots.
The datacenters would be built next to wind farms if it wasn’t for oil industry bribes. It’s just a scape goat.
Gas turbines need to be outlawed outright and we need proper regulation for all industries. They focus on getting everyone pissed about AI so you don’t start talking about hanging oil execs.
Ellendale ND has a DC in progress. The region has a power surplus. What you are wishing for is already started.
Lmao, price of fuel right now has people going “por que no los dos?”
You can run model locally, a gaming PC is enough, may-be not for cuting edge models but if you want to generate clues for a RPG or rephrase a letter it’s good enough. You can look for LM studio and Stability matrix for example
Bonus if you’re using solar power to generate your responses.
It is absolutely possible to run 100% locally, but in practice, at the high end, only quantized models. The full size top end models require hundreds of GB of video ram, and while you can buy that, it’s stupidly expensive. Quantized models can often perform nearly as well with a small fraction of the ram, but they do sacrifice a little in precision.
These models (well the good ones) ultimately all trace their origins to what you’d likely consider “stolen” data. Whether that’s ethical is debatable. If you’re in the “information should be free” camp, there may not be an issue here.
As for the power/environmental impact, for what they do LLMs are actually very low impact per-request. If you’re concerned about your personal AI power use, then I hope you never fly in an airplane, and minimize your driving because those are much bigger issues.
It’s the scale of use that makes AI an environmental problem, and that’s a question about corporate use of AI, not personal use of AI.
The power impact was something that in the early days worried me. Looking into it I agree with you to some degree. Like using it instead of a search engine I think is by and large a wash. One prompt will likely take more energy but will give you information that likely would have required searching several times modifying the words and jumping between sites which are rendering all sorts of things. Heck If I booted into a command line and connected to an llm Im almost sure it would be significantly less energy. If you chat for entertainment instead of streaming vidoe also lower energy use. Now I think one thing is in making things. It lets people who otherwise couldn’t make pictures and videos and code. In the large majority of cases what is made is going to be disposed even for folks that eventually make something they care to keep around or use. While using software to do the same uses a lot of energy the only people doing it generally where making long lasting things for projects or such. So that is where I question it. Still I will have it make a picture to use in an rpg or such.
From that perspective local LLMs sound more like classical piracy. Not ethical, not FOSS, but out of the hands of greedy corporates.
Some of the open models claim not to use pirated content. I don’t necessarily believe it, bit its being claimed.
they are also doing distillation from the big models from OpenAI & Claude, so open-weight models with similar capability can be available for free and also reduce the big AI companies’ ability to profit from stolen data.
You need some decent hardware to do it, but yes one can absolutely run capable local models. The 7-8B models run fast and are probably decent for light coding, but they make piss poor chat bots. 27-36B models get much better outputs but take more hardware, and still not great in a chat interface (but far better than 8B). They run slow enough on my MacBook that it’s just a toy for fun, despite unified memory. I aspire to having a local AI server then can run models in the 100B-350B range. It’s a pricy bit of hardware, but I can leverage it to do all the things one might want to do but avoid for data privacy purposes. I try to avoid oversharing, but when I look at what ChatGPT knows about me, I’m giving way more away than I want to. I look forward to the day I don’t have to use it at all.
Ethically is a slightly different story; there are a number of claims people like to make:
Plagiarism: not in any significant sense. Like any of us might borrow a turn of phrase, so might AI, but any significant body of text — even a paragraph — is very unlikely to match exactly to an ingested work. That’s why they need to ingest so much different source material — so no one author or text dominates the output.
Stealing: Mostly yes. I think there is a significant difference between ingesting all of Reddit (where people including myself were sharing their words with the world) and feeding in copyrighted works of published authors without permission. I think when you turn around and sell access to a model built on material you didn’t pay for, that’s wrong. I do think open models available to everyone are far less evil because they take in information they don’t own and give back a tool they don’t own, having contributed significant money in the form of training.
Infrastructure: Yes but… that’s how everyone does it from Walmart to sports stadiums. Don’t single out data centers when the entire system is corrupt. Companies should pay for transmission lines and capacity they expect to draw. All companies, not just AI.
Environment: Yes but… the amount of water used sounds high on a global basis, but it’s utterly dwarfed by agriculture and some other industries. Now some say, yeah but we have to eat, but do we have to eat almonds specifically? Beef? Do we have to grow food in deserts? Data center water usage looks big in raw numbers, but it’s not in percentage terms. That being said, it can be a huge problem locally where a community might have a sustainable water supply but a big data center tips the scales.
The latter two are IT industry issues that are driven by AI but aren’t going away when and if the bubble pops. Imagine the privacy nightmare that is coming. In fact, I’m not convinced that AI isn’t just a cover for unprecedented data collection.
I’ll take a crack at it.
I run AI on an old Nvidia p40 purchased online for less than $200. The models I run on it came predominantly from Chinese companies and groups. On balance, they:
-
use much less fossil fuel in producing these LLMs than their Western counterparts
-
produce models that run much better on lower end consumer hardware
As for training data, the inputs used to create these models, it varies greatly but the most popular line, Qwen, Came from the company’s own data from operating such huge networks and systems for so long
I really fail to see how any of that is worse than playing a video game.
None of this takes away from the very real issues around data center build out and Western companies using the systems to scare people and continue an economic bubble. That’s all true and bad. But there’s nuance. Not all AI is created equal
use much less fossil fuel in producing these LLMs than their Western counterparts
Utterly false, given that all the decent Chinese models (especially Qwen) are distilled from Western frontier models.
They literally could not exist without the enormously carbon emissive western models existing first to train them.
yeah the comment was bad enough I looked at the user. been around for 2 years and no posts or comments till this one. made a note on the account.
-
They are all trained on copyrighted material without permission, no LLM is ethical
What’s not ethical is a copyright system thag enforces artificial scarcity where there is no need for it.
Piracy is not stealing, and is not inherently unethical.
Everything is copyrighted regardless of it being owned by a corporation or blogger
I don’t partially care about the former
Human brains are all trained on copyrighted material without permission, no human brain is ethical
do you consider “piracy” like zlibrary and annas archive unethical?
If a business is using them, yes.
You can also train your own local models with license free material if you wish! I think one of the easiest ways to get into that is by using software like unsloth (that’s the one i am using), an open source no-code tool which can be both used to train models on whatever data you wish and to run models either locally or using an inference provider.
Quick example for something like that which is also not dependent on copyrighted material is RAG, where you can provide the 400-page manual for something and then can chat with “the document” to get explanations, ask quick questions without searching for possibly multiple occurrences of a specific term and similar stuff.
Doesn’t RAG require a pretrained model still? Presumably on copyright material?
It does not per se need a model using copyrighted stuff. You can get by using a model without that - basic language skills and technical jargon can easily be trained with open material.
In theory that makes sense, but does this actually exist?
Some don’t speak any “language” they are trained to “speak” and “think” in terms of election orbitals and bonding energy. They are used in pharma and materials science to work on intractable problems like superconductivity and meds for Parkinson’s.
that’s not really an LLM then, is it?
It is, cause it uses the same architecture, https://www.geeksforgeeks.org/artificial-intelligence/exploring-the-technical-architecture-behind-large-language-models/
an example: https://github.com/bowang-lab/scGPT “talks” in RNA and is incredibly accurate and able to do predictions
Some companies build out more electrical capacity than they use. Substations are funded by the DC and built to twice the desired capacity - and half is delivered to the community. The generators are run only when utility power goes down, and run on diesel.
However, not all companies are this responsible and the grievously irresponsible ones get the press and make the responsible companies look bad.
Can you name how many data centers have actually built out excess capacity before they started operating?
Most of the ones I’ve heard have simply claimed that they might eventually, which is a corporate way of saying we’ll walk that back once we think no one is paying attention.
I run models locally on my 5070 i bought a year back for around 550€ (now costs around 900€ - insanity). Yes, it runs completely locally with pretty good results, and i am using an AM4 platform, which is now a few years old, and was also able to run smaller models with my 3070Ti with good speed.
My viewpoint regarding the models themselves is that since everything on the web has been used including everything i have put on the web in the last decades, there is also not much of an ethical problem here when using these things in a non-commercial setting - i am not making any money with it and i am not disseminating the output, and therefor i also do not cut into the profits of any creator (i mainly use it to automate tedious tasks or to get a complicated RegEx/sed/awk situation under control without breaking my brain - no cultural output like text or images). On a larger scale i would prefer it if models and datasets were under UNESCO stewardship - free for non-commercial use, with licenses being sold to corporations and organizations, where the income from the licenses get distributed to the people providing for the datasets (in this case in an opt-out basis - these models already represent a cultural snapshot of humanity, and the assumption is that artists WANT to add to human culture) or for financing young artists.
They call it open-weighted and not open-source because they don’t have the training material (stolen books lol) but still want to pretend they are l33t hackers not bound to BigTech.
Yes, it’s possible to use it locally with a $2000 computer (that’s for the cheapest ones) but you’ll only get a few words per second. It doesn’t matter to the vibe coders who don’t know how to code.
The software to run that is open-source but the training of the model (the “open-weight” black box) requires to destroy the environment at least once.
Last but not least the free models are obviously censored but people don’t care about censorship anymore for some reason. “Tiananmen didn’t happen? Not my problem” without understanding that more is hidden.
Anyway, no, there is nothing ethical about it.
a) Yes, today it costs 2000$ because of the cost explosion - my setup cost me around 1400$, and my GPU was already bought when prices were rising. No, i do not get a few words per second, i get around 50-60 token/s, which is more than enough for personal use. (Edit: and that is WITH CPU offloading, where layers that don’t fit into VRAM get placed into system RAM, and i am still running DDR4 to boot)
b) If the completely insane AI corpos would stop training humongous models to chase after non-achievable AGI, the one-time investments would have paid off by now for local use. This situation has nothing to do with local models but insane billionaires and execs.
c) go ahead and lookup abliterated (not a typo) models on HuggingFace - these do NOT refuse any requests, because that has been pruned out. These answer everything about Tiananmen, DEI topics like erosion of LGBT and womens rights in the west and whatever atrocities any group might have commited, while also telling you about whatever you want to know. I run only these models, because i refuse to partake in censorship.
Edit: Regarding the training material: This training material also consists of MY output over the decades on the web. Therefor, i do not have qualms using these models for non-commercial usage without disseminating the output - classic personal usage, mainly for automating tasks that are not easily done by hand such as grabbing game descriptions from steam, condensing them down to a short sentence and putting the result in a database, together with the corresponding steam tags, and automatically searching the web for information about the game when there is no steam page for it - which is an insane amount of work to do by hand for my library of more than 30k titles. I let it work during off hours on that task.








