Wikipedia Grapples With Chatbots: Should It Allow Their Use For Articles? Should It Allow Them To Train On Wikipedia?
from the questions,-questions dept
There have been various chapters in the new large language models (LLMs) story. First, people were amazed that systems like ChatGPT could write a sonnet about bananas in the style of Shakespeare, and in just a few seconds. Soon, though, they realized that chatbots’ replies might be grammatically correct, but they were frequently peppered with false information that the system simply made up, often with equally fake references. We’re now at the stage where many are starting to think through the deeper implications of using LLMs, with all their powers and flaws, and how they will affect current working (and living) practices. As a post on the Vice site explains, one group grappling with this issue is the Wikipedia community:
During a recent community call, it became apparent that there is a community split over whether or not to use large language models to generate content. While some people expressed that tools like Open AI’s ChatGPT could help with generating and summarizing articles, others remained wary.
Wikipedia already has a draft policy on how LLMs can be used when writing Wikipedia entries. The draft provides an excellent summary of some of the key problems of using chatbots, many of which will be faced by people in other domains. Here are the main points from the “basic guidance” section:
Do not publish content on Wikipedia obtained by asking LLMs to write original content or generate references. Even if such content has been heavily edited, seek other alternatives that don’t use machine-generated content.
You may use LLMs as a writing advisor, i.e. asking for outlines, asking how to improve paragraphs, asking for criticism of text, etc. However, you should be aware that the information they give to you can be unreliable and flat out wrong. Use due diligence and common sense when choosing whether to incorporate the LLM’s suggestions or not.
You may use LLMs for copyediting, summarization, and paraphrasing, but note that they may not properly detect grammatical errors or keep key information intact. Use due diligence and heavily edit the response. Don’t hesitate to ask the LLM to correct deficiencies such as missing information in a summary or an unencyclopedic, e.g. promotional tone.
You are responsible for making sure that using an LLM will not be disruptive to Wikipedia.
You must denote that a LLM was used in the edit summary.
LLM-created works are not reliable sources. Unless their outputs were published by reliable outlets with rigorous oversight, they should not be cited in our articles.
It would be foolish to try to forbid Wikipedia contributors from using chatbots to help write articles: people would use them anyway, but would try to hide the fact. A ban would also be counterproductive. LLMs are simply tools, just like computers, and the real issue is not whether to use them, but how to use them properly. The guidelines listed above essentially amount to “yes, you can use chatbots to help you write and improve your writing, but they should not be relied upon unquestioningly.” That means human input and checking afterwards are indispensable. Also important is flagging up that LLMs were used in some way, so that users of Wikipedia know where information is coming from, and can be alert to possible problems arising from this fact.
The Wikipedia draft policy concentrates on how LLMs’ output might be used to create material for Wikipedia entries. The Vice article points out that there is another question, about whether there should be restrictions on how LLMs can use Wikipedia entries as part of the machine learning process:
The [Wikipedia] community is also divided on whether large language models should be allowed to train on Wikipedia content. While open access is a cornerstone of Wikipedia’s design principles, some worry the unrestricted scraping of internet data allows AI companies like OpenAI to exploit the open web to create closed commercial datasets for their models. This is especially a problem if the Wikipedia content itself is AI-generated, creating a feedback loop of potentially biased information, if left unchecked.
That concern seems overblown. Low-quality training materials can cause chatbots to produce questionable or downright harmful outputs. An obvious way to counter that would be to encourage the use of high-quality input that has undergone some kind of fact checking. Wikipedia is one of the best and largest sources of such material, and in hundreds of languages. Provided the final Wikipedia policy on LLMs requires human checks on chatbot output, as proposed in the draft, the use of Wikipedia articles for training LLMs should surely be encouraged with the aim of making chatbots better for everyone.
Follow me @glynmoody on Mastodon.
Filed Under: chatbots, chatgpt, llms, wikipedia
Companies: openai


Comments on “Wikipedia Grapples With Chatbots: Should It Allow Their Use For Articles? Should It Allow Them To Train On Wikipedia?”
The real danger of training ML on human output… is that human output contaminates the training model.
You risk getting stuff like this: https://www.nuklearpower.com/2008/05/31/episode-999-like-a-hobby/
(the behavior describe, not the excellent humor… by a human ….)
I don’t think Bard and Chat-GPT are evil. But they are products intended to reduce the amount of time it takes people to find things on search engines, which is only sort of related to finding the truth.
For an encyclopedia, which exists to provide truthful information, to rely on chatbots, the chatbots must draw only from reliable sources, sift out junk sources, measure the relative value of sources, and understand how those sources fit into the broader context of the article. Right now, no chatbot exists that can do these things. So a ban, or at least a strict human review policy, on using chatbots to write articles should be on the table, even if some bad faith actors will just ignore it.
So I echo Wikipedia’s concern over bots using their own writing as training. It would quickly result in a thousand Jar-Edo Wenses — the same self-referential wrongness that got us on the eight-spiders-per-year diet.
If chatbots don’t generate Wikipedia content, though, I don’t see much issue with them scraping Wikipedia for training, even if it does inure to the benefit of massive tech corporations. Worst case, they become better products.
Re:
Only if the Wikimedia Foundation agrees to it, be it through a contract or something else.
Otherwise, it’d be rather unfair.
Re: Re:
You know, the more I think about it, the more I agree that LLM companies really ought to get permission or at least include attribution whenever they use Wikipedia content, since LLMs have always served a commercial purpose. Which would also comply with Wikipedia’s CC BY-SA 3.0 content license.
The share-alike aspect of the CC BY-SA 3.0 license would bother LLM makers the most I think. But they’d just have to live with it.
Re: Re: Re:
They should, but one of the main reasons why we’re even having this damn conversation, and why Wikipedia is that hesitant to allow LLM companies to use wikipedia as a training source, is because these companies don’t give a fuck about permissions.
Remember: NovelAI trained its LLM on danbooru. ChatGPT used its research arm to get a training dataset. I believe one other company is also using a dataset from its research arm as well.
Re: Re: Re:2
Okay, copyright maximalist.
“You are free until you do something we don’t like”?
The same thing happens with some GPL or open source computer programs, as their creators talk a good game about freedom but then try to impose restrictions when someone they don’t like uses their code.
People have contributed to Wikipedia under the condition that their work be freely redistributable and copyable by anyone who wants it. The would be masters of Wikipedia don’t get to limit that because they want to impose their fear on everyone else.
Re:
The creators of the GPL have flat out come out against using copyright to restrict who can use their code based on actions that are unrelated to the license itself… so that statement kind of fails to be true.
Re: Re:
It does not fail to be true. Many people who are not the creators of the GPL license their code under it. They may have their own different views. Something similar happened when Wizards of the Coast proposed retroactive modifications to the open license they had granted for derivative Dungeons & Dragons works.
Note that beneficent motives should not matter. The AI fearmongers may believe in their doomsday scenarios. WotC was interested in avoiding the spread of old racist and sexist tropes that were part of D&D. But freedom should trump everything.
Re: Re: Re:
It’s one thing to desire transparency. And this is usually a good thing.
It’s another to let violent, treasonous white supremacists to run the show.
We do not need to tolerlate the intolerant.
Re: Re: Re:
It’s not similar at all. Wizards of the Coast made a new license. (Yes, a new version of a license is a new license.) Revoking a GNU GPL (which is perpetual and implicitly irrevocable) and replacing it with a license with fewer freedoms would be similar, but that’s not nearly the same as nominally keeping the GPL and adding restrictions outside of the license. Also, the latter is explicitly prohibited by the GPLs.
Re:
Hey, if you want to donate your time and expertise to the wealthiest companies in the world, fill your boots. Getting huffy about freedom doesn’t make you any less of a mark though.
Re:
Contributors to Wikipedia grant the license to reuse their contributions only under the “share-alike” condition, similar to the copyleft condition embodied in GPL. The licenses are indeed for anyone (so if some contributors object to any use by someone they don’t like, they are indeed in the wrong), but not for any use that violates the conditions that are there to guarantee this freedom to the users of the derivative works (so complaints about such uses are just stating that the use violates the original conditions of the contribution).
Quoting from Wikipedia terms, where they state their licensing conditions, emphasis added:
It may be different when someone first says anyone can do anything with their contribution and then goes on to complain when someone does something they don’t like. But Wikipedia does not say that in the first place.
Re: Re:
Cue pointless bickering over whether an LLM is a derivative work…
Re: The FSF
I believe that the FSF (free software foundation) was soliciting feedback with regards to whether they wanted to sue OpenAI for the use of CoPilot not reproducing the GPL license.
There is a lawsuit by a John Doe litigant, who also sought damages in addition to injunctive relief, and the damages were dismissed with prejudice with leave, and injunctive relief was not dismissed but held for summary judgement, and issues relating to whether the litigant even has jurisdictional standing as a john doe litigant who has not actually claimed or registered any specific copyrighted code.
Re:
You should be more specific about what you mean by “impose restrictions” and “when someone they don’t like uses their code”. All you have right now is a vague generalization. And regardless, the license violations you theoretically refer to don’t detract from merits of using GPL or other libre licenses in a valid way.
Tangent: I would like copyleft licenses (such as the GNU GPLs and CC BY-SA 4.0) to apply to AI outputs (because it results in more freedom, in a manner of speaking), but fair use might make licenses on AI inputs inapplicable to the corresponding outputs.
Use chatbots if you want something hilarious or just poor.
The problem with scraping is that most humans producing output which can be scraped are as bad as, or worse than chatbot when it comes to accuracy. So yeah, that training data…
ChatGPT (and other LLMs) produce text which “sounds like it would follow”. If asked to document its sources, it will invent plausible-seeming footnotes, even URLs. Doesn’t mean that any of it is true. Just that it sounds plausible.
I think it should be banned for a simple reason — it would be the digital LLM equivalent of incest.
LLM-assisted text which is inserted into a knowledge repository that is supposed to be factual and serves as a training ground for future LLMs would basically result in a very long and drawn out process of the LLM training itself on its own writing, as LLM-assisted text begins to become on par/outweigh the proportion of human writers on the platform – for the same reason you wouldn’t wanna train an LLM on itself, I think Wikipedia should ban it outright. Even if you stick to simply fixing up syntax and grammar, it’ll result in a distinct LLM-dialect in a few years.
Re:
The ouroboros comes to mind, yeah.
Addressing the technical concern raised in the last blockquote while completely ignoring the moral objection is so perfectly Techdirt I might have to get it framed.