The Sunday Read: ‘Wikipedia’s Moment of Truth’
hi I'm John gertner I'm a contributor to. the New York Times magazine and I write. about science and technology. this week's Sunday read is a story I. wrote for the magazine about Wikipedia. a story that explains how the 22 year. old wonky online encyclopedia we've all. consulted at one point is so Central to. building artificial intelligence right. now. so over the last few years computer.
scientists have been creating what are. known as large language models which are. the AI brains the power the chat Bots. like chat TPT. and in order to build a large language. model they needed to gather vast. knowledge Banks of information. and I mean it's sort of dizzying how. much information we're talking about. here some models ingest upwards of a. trillion words. and it all comes from public sources. like Wikipedia or Reddit or Google's.
patent database. what makes Wikipedia special is not just. that it's free and accessible but also. that it's very highly formatted. it contains just a tremendous amount of. factual information that's maintained by. a community of about forty thousand. active editors in the English language. version alone. the problem with these new AI chat Bots. is that their fundamental goal is to. converse with the user with a kind of.
human fluency of language but they're. not built to regurgitate data or to. really be precise. so whether you're trying to understand. historical topics or political upheavals. or pandemics these Bots greatly simplify. the world in a way that's maybe not. conducive at all to our best interests. as human beings. AI chatbots have even been known to. hallucinate and conjure falsehoods from.
Whole cloth and another problem is that. if they're fed only on their own. synthetic data these systems essentially. break down. so if we were to go to AI instead of. Wikipedia to find information to solve. our problems to answer questions. what would happen in the future where. our knowledge is factually unreliable. as I recorded this story I read a lot of. what are called Community notes which.
are the logs of Wikipedia editor. meetings that they transcribe and make. public. and in one recent meeting editors shared. their worries about AI. what's it going to do to Wikipedia. I remember reading the notes for this. meeting and one line from an editor. popped out at me. we want a future where knowledge is. created by humans. and I thought well that's really the.
essence of it isn't it. can we really choose at this point the. future we want. so here's my article Wikipedia's Moment. of Truth. read by Brian yishun. in early 2021 a Wikipedia editor peered.
into the future and saw what looked like. a funnel cloud on the horizon the rise. of gpt3 a precursor to the new chat Bots. from openai. when this editor a prolific wikipedian. who goes by the handle barkeep 49 on the. site gave the new technology a try he. could see that it was untrustworthy the. bot would readily mix fictional elements. a false name a false academic citation.
into otherwise factual and coherent. answers but he had no doubts about its. potential I think ai's day of writing a. high quality encyclopedia is coming. sooner rather than later he wrote in. death of Wikipedia an essay that he. posted under his handle on Wikipedia. itself. he speculated that a computerized model. could in time displace his beloved. website and its human editors just as.
Wikipedia had supplanted the. Encyclopedia Britannica which in 2012. announced it was discontinuing its print. publication. recently when I asked this editor he. asked me to withhold his name because. Wikipedia editors can be the targets of. abuse. if he's still worried about his. encyclopedia's fate. he told me that the newer versions made.
him more convinced that chat GPT was a. threat. it wouldn't surprise me if things are. fine for the next three years he said of. Wikipedia and then all of a sudden in. year four or five things drop off a. cliff. Wikipedia marked its 22nd anniversary in. January it remains in many ways a. throwback to the internet's utopian. early days when experiments with open. collaboration anyone can write and edit.
for Wikipedia had yet to see the digital. terrain to multi-billion Dollar. corporations and data miners advertising. schemers and social media propagandists. the goal of Wikipedia as its co-founder. Jimmy Wales described it in 2004 was to. create a world in which every single. person on the planet is given free. access to the sum of all human knowledge. the following year Wales also stated we.
helped the internet not suck. Wikipedia now has versions and 334. languages and a total of more than 61. million articles. it consistently ranks among the world's. 10 most visited websites yet is alone. among that select group whose usual. leaders are Google YouTube and Facebook. in issuing the profit motive. Wikipedia does not run ads except when. it seeks donations and its contributors.
who make about 345 edits per minute on. the site are not paid. in seeming to repudiate capitalism's. imperatives its success can seem. surprising even mystifying. some wikipedians remark that their. Endeavor Works in practice but not in. theory. Wikipedia is no longer an encyclopedia. or at least not only an encyclopedia. over the past decade it has become a. kind of factual netting that holds the.
whole digital world together. the answers we get from searches on. Google and Bing or from Siri and Alexa. how old is Joe Biden or what is an ocean. submersible derive in part from. Wikipedia's data having been ingested. into their knowledge Banks. YouTube has also drawn on Wikipedia to. counter misinformation. the new AI chatbots have typically. swallowed Wikipedia's Corpus too.
embedded deep within their responses to. queries is Wikipedia data and Wikipedia. text knowledge that has been compiled. over years of painstaking work by human. contributors. while estimates of its influence can. vary Wikipedia is probably the most. important single Source in the training. of AI models. without Wikipedia generative AI wouldn't. exist says Nicholas Vincent who will be. joining The Faculty of Simon Fraser.
University in British Columbia this. month and who has studied how Wikipedia. helps support Google searches and other. information businesses yet as Bots like. chat GPT become increasingly popular and. sophisticated Vincent and some of his. colleagues wonder what will happen if. Wikipedia outflanked by AI that has. cannibalized It suffers from disuse and. dereliction. in such a future a death of Wikipedia.
outcome is perhaps not so far-fetched. a computer intelligence it might not. need to be as good as Wikipedia merely. good enough is plugged into the web and. seizes the opportunity to summarize. Source materials and news articles. instantly the way humans now do with. argument and deliberation. on a conference call in March that. focused on ai's threats to Wikipedia as. well as the potential benefits the.
editor's hopes contended with anxiety. while some participants seemed confident. that generative AI tools would soon help. expand Wikipedia's articles and Global. reach others worried about whether users. would increasingly choose chat gbt fast. fluent seemingly oracular over a wonky. entry from Wikipedia. a main concern among the editors was how. wikipedians could defend themselves from.
such a threatening technological. interloper and some worried about. whether the digital realm had reached. the point where their own organization. especially in its striving for accuracy. and truthfulness was being threatened by. a type of intelligence that was both. factually unreliable and hard to contain. one conclusion from the conference call. was clear enough we want a world in. which knowledge is created by humans. but is it already too late for that.
back in 2017 the Wikimedia foundation. and its community of volunteers began. exploring how the encyclopedia and its. sister sites like wikidata and Wikimedia. Commons with their offerings of free. information and images could evolve by. the year 2030. the plan was to ensure that the. foundation the non-profit that oversees. Wikipedia could protect and share the.
world's information in perpetuity. one outcome of that 2017 effort which. included a Year's worth of meetings was. a prediction that Wikimedia would become. the essential infrastructure of the. ecosystem of free knowledge. another conclusion was that trends like. online misinformation would soon require. far more vigilance. and a research paper commissioned by the. foundation found that artificial. intelligence was improving at a rate.
that could change the way that knowledge. is gathered assembled and synthesized. for that reason the rollout of chat gbt. did not elicit surprise inside the. Wikipedia Community though several. editors told me they were shocked by the. speed of its adoption which needed just. two months after its release in late. 2022 to gain an estimated 100 million. users. despite its stodgy appearance Wikipedia.
is more tech savvy than casual users. might assume with a small group of. volunteers to oversee millions of. Articles it has long been necessary for. highly experienced editors often known. as administrators to use semi-automated. software to identify misspellings and. catch certain forms of intentional. misinformation. and because of its open source ethos the. organization has at times Incorporated. technology made freely available by tech.
companies or academics rather than go. through a lengthy and expensive. development process on its own. we've had artificial intelligence tools. and Bots since 2002 and we've had a team. dedicated to machine learning since. 2017. Selena deckleman wikimedia's Chief. technology officer told me they're. extremely valuable for semi-automated. Content review and especially for. translations. [Music]. how Wikipedia uses Bots and how Bots use.
Wikipedia are extremely different. however. for years it has been clear that. fledgling AI systems were being trained. on the site's articles as part of the. process whereby Engineers scrape the web. to create enormous data sets for that. purpose. in the early days of these models about. a decade ago Wikipedia represented a. large percentage of the scraped data. used to train machines.
the encyclopedia was crucial not only. because it's free and accessible but. also because it contains a motherload of. facts and so much of its material is. consistently formatted. in more recent years as so-called large. language models or llms increased in. size and functionality these are the. models that power chat Bots like chat. GPT and Google sparred they began to. take in Far larger amounts of.
information. in some cases their meals added up to. well over a trillion words. the sources included not just Wikipedia. but also Google's patent database. government documents reddit's q a corpus. books from online libraries and vast. numbers of news articles on the web. but while Wikipedia's contribution in. terms of overall volume is shrinking and. even as tech companies have stopped.
disclosing what data sets go into their. AI models it remains one of the largest. single sources for llms. Jesse Dodge a computer scientist at the. Allen Institute for AI in Seattle told. me that Wikipedia might now make up. between three and five percent of the. scrape data and llm uses for its. training. Wikipedia going forward will forever be. super valuable Dodge points out because.
it's one of the largest well-curated. data sets out there. there is generally a link he adds. between the quality of data a model. trains on and the accuracy and coherence. of its responses. in this light Wikipedia might be seen as. a sheep caught in the jaws of a wolfish. technology Marketplace. a free site created an achingly good. faith sharing knowledge is by nature and.
act of kindness Wikimedia noted in 2017. on a page devoted to its strategic. direction is being devoured by companies. whose objectives like charging for. subscriptions as open AI recently began. doing for its latest model don't jive. with its own. yet the relationships are more. complicated than they appear. Wikipedia's fundamental goal is to. spread knowledge as broadly and freely. as possible by whatever means.
about 10 years ago when site. administrators focused on how Google was. using Wikipedia they were in a situation. that pressaged the Advent of AI chatbots. Google's search engine was able at the. top of its query results to present. wikipedians work to users all over the. world giving the encyclopedia far. greater reach than before an apparent. virtue. in 2017 three academic computer.
scientists Conor McMahon Isaac Johnson. and Brent Hecht conducted an experiment. that tested how random users would react. if just part of the contributions made. to Google's search results by Wikipedia. were removed. the academics perceived an extensive. interdependence. Wikipedia makes Google a significantly. better search engine for many queries. and Wikipedia in turn gets most of its.
traffic from Google. one up shot from the collision with. Google and others who repurpose. Wikipedia's content was the creation two. years ago of Wikimedia Enterprise a. separate business unit that sells access. to a series of application programming. interfaces that provide accelerated. updates to Wikipedia articles. depending on whom you ask the Enterprise. unit is either a more formalized way for. tech companies to direct the equivalent.
of large charitable donations to. Wikipedia Google Now subscribes and. altogether the unit took in 3.1 million. dollars in 2022 or a way for Wikipedia. to recoup some of the financial value it. creates for the digital world and thus. help fund its future operations. practically speaking Wikipedia's. openness allows any tech company to. access Wikipedia at any time but the. apis make new Wikipedia entries almost.
instantly readable. this speeds up what was already a pretty. fast connection. Andrew Lee a consultant who works with. museums to put data about their. collections on Wikipedia told me he. conducted an experiment in 2019 to see. how long it would take for a new. Wikipedia article about a pioneering. balloonist named Vera Simons to show up. in Google search results. he found the elapsed time was about 15.
minutes. still the close relationship between. search engines and Wikipedia has raised. some existential questions for the. latter. ask Google what is the Russia Ukrainian. war and Wikipedia is credited with some. of its material briefly summarized. but what if that makes you less likely. to visit Wikipedia's article which runs. to some 10 000 words and contains more. than 400 footnotes from the point of.
view of some of Wikipedia's editors. reduced traffic will oversimplify our. understanding of the world and make it. difficult to recruit a new generation of. contributors it may also translate into. fewer donations in the 2017 paper the. researchers noted that visits to. Wikipedia had indeed begun to decline. and the phenomenon they identified. became known as the Paradox of reuse.
the more Wikipedia's articles were. disseminated through other outlets and. media the more imperiled was Wikipedia's. own health. with AI this reuse problem threatens to. become far more pervasive. Aaron hafaker who led the machine. learning research team at the Wikimedia. foundation for several years and who now. works for Microsoft told me that search. engine summaries at least offer users. links and citations and a way to click.
back to Wikipedia. the responses from large language models. can resemble an information Smoothie. that goes down easy but contains. mysterious ingredients. the ability to generate an answer has. fundamentally shifted he says noting. that in a chat GPT answer there is. literally no citation and no grounding. in the literature as to where that. information came from. he contrasts it with the Google or Bing. search engines this is different this is.
way more powerful than what we had. before almost certainly that makes AI. both more difficult to contend with and. potentially more harmful at least from. Wikipedia's perspective. a computer scientist who works in the AI. industry but is not permitted to speak. publicly about his work told me that. these Technologies are highly. self-destructive threatening to. obliterate the very content which they. depend upon for training it's just that.
many people including some in the tech. industry haven't yet realized the. implications. [Music]. Wikipedia's most devoted supporters will. readily acknowledge that it has plenty. of flaws. the Wikimedia Foundation estimates that.
its English language site has about 40. 000 active editors meaning they make at. least five edits a month to the. encyclopedia. according to recent data from the. Wikimedia Foundation about 80 percent of. that cohort is male and about 75 percent. of those from the United States are. white which has led to some gender and. racial gaps in Wikipedia's coverage and. lingering doubts about reliability. remain.
for a popular article that might have. thousands of contributors Wikipedia is. literally the most accurate form of. information ever created by humans Amy. Brockman a professor at the Georgia. Institute of Technology told me. but Wikipedia's short articles can. sometimes be hit or miss. they could be total garbage says. Brookman who is the author of the recent. book should you believe Wikipedia. an erroneous fact on a rarely visited. page may endure for months or years.
and there continues to exist the. ever-present threat of vandalism or. tampering with an article. in 2017 for instance a photo of the. Speaker of the House Paul Ryan was added. to the entry on invertebrates. as a Wikipedia editor whose first name. is Jade put it to me we have a number of. I would say almost professional trolls. who must dedicate just about as much. time to creating spam creating vandalism.
harassing people as I dedicate to. improving Wikipedia. several academics told me that whatever. Wikipedia's shortcomings they view the. encyclopedia as a consensus truth as one. of them put it it acts as a reality. check in a society where facts are. increasingly contested. the truth is less about data points how. old is Joe Biden than about Complex. events like the kova 19 pandemic in.
which facts are constantly evolving. frequently distorted and furiously. debated. the truthfulness quotient is raised by. Wikipedia's transparency. most Wikipedia entries include footnotes. links to Source materials and lists of. previous edits and editors and. experienced editors are willing to. intercede when an article appears. incomplete or lacks what wikipedians. call verifiability.
moreover Wikipedia's guidelines insist. that its editors maintain an npov. neutral point of view or risk being. overruled or in the Argo of Wiki culture. reverted. and the site has a bent towards. self-examination you can find long. disquisitions on Wikipedia that explore. Wikipedia's own reliability. an entry on how Wikipedia has fallen. victim to hoaxes runs to more than 60.
printed pages. as difficult as The Pursuit Of Truth can. be for wikipedians though it seems. significantly harder for AI chatbots. chat gbt has become infamous for. generating fictional data points or. false citations known as hallucinations. perhaps more Insidious is the tendency. of bots to oversimplify complex issues. like the origins of the ukraine-russia. war for example. one worry about generative AI at.
Wikipedia whose articles on medical. diagnoses and treatments are heavily. visited is related to health information. a summary of the March conference call. captures the issue. we're putting people's lives in the. hands of this technology for example. people might ask this technology for. medical advice it may be wrong and. people will die. this apprehension extends not just to. chatbots but also to new search engines.
connected to AI Technologies in April a. team of Stanford University scientists. evaluated four engines powered by AI. Bing chat Niva AI perplexity Ai and. uchat and found that only about half of. the sentences generated by the search. engines in response to a query could be. fully supported by factual citations. we believe that these results are. concerningly low for systems that may.
serve as a primary tool for information. seeking users the researchers concluded. especially given their facade of. trustworthiness. [Music]. what makes the goal of accuracy so. vexing for chat Bots is that they. operate probabilistically when choosing. the next word in a sentence. they aren't trying to find the light of. Truth in a murky world.
these models are built to generate text. that sounds like what a person would say. that's the key thing Jesse Dodge says so. they're definitely not built to be. truthful. I asked Margaret Mitchell a computer. scientist who studied the ethics of AI. at Google whether factuality should have. been a more fundamental priority for AI. Mitchell who says she was fired from the. company after criticizing the direction. of its work Google says she was fired.
for violating the company's security. policies said that most would find that. logical. this Common Sense thing shouldn't we. work on making it factual if we're. putting it forward for fact-based. applications well I think for most. people who are not in Tech it's like why. is this even a question. but Mitchell said the priority is that. the big companies now in frenzied. competition with one another are. concerned with introducing AI products.
rather than reliability. the road ahead will almost certainly. lead to improvements Mitchell told me. that she foresees AI companies making. gains in accuracy and reducing biased. Answers by using better data. the state of the art until now has just. been a laissez-faire data approach she. said you just throw everything in and. you're operating with a mindset where. the more data you have the more accurate.
your system will be as opposed to the. higher quality of data you have the more. accurate your system will be. Jesse Dodge for his part points to an. idea known as retrieval whereby a. chatbot will essentially consult a high. quality Source on the web to fact check. and answer in real time. it would even cite precise links as some. ai-powered search engines Now do. without that retrieval element Dodge. says I don't think there's a way to.
solve the hallucination problem. otherwise he says he doubts that a. chatbot answer can gain factual parity. with Wikipedia or the Encyclopedia. Britannica. Market competition might help prompt. improvement too Owen Evans a researcher. at a non-profit in Berkeley California. who studies truthfulness in AI systems. pointed out to me that openai now has. several Partnerships with businesses and.
those firms will care greatly about. responses achieving a high level of. accuracy Google meanwhile is developing. AI systems to work closely with medical. professionals on disease detection and. Diagnostics. there's just going to be a very high bar. there he adds so I think there are. incentives for the companies to really. improve this. at least for now ai companies are. focusing on what they call fine tuning.
when it comes to factuality. sandini argoal and Girish sastri. researchers at open AI the company that. created chat GPT told me that their. newer AI model gpt4 has made significant. improvements over earlier models in what. they called factual content. those advances stem mainly from a. process known as reinforcement learning. with human feedback to help AI models. differentiate between good and bad.
answers. but chat GPT clearly has a way to go. both to fix hallucinations and to. provide complex multi-layered and. accurate answers to historical questions. when I asked argoal whether openai. systems could ever be completely. accurate or offer 400 footnotes she said. that it was possible. but there might always exist a tension. between a model's ambition to be factual.
and its efforts to be creative and. fluent. as an AI developer she explained the. goal was not for a chat model to. regurgitate data it had been trained on. rather it was to see patterns of. knowledge it could relate to users in. fresh conversational language. in the future sastri added AI systems. might interpret whether a query requires. a rigorous factual answer or something. more creative.
in other words if you wanted an. analytical Report with citations and. detailed attributions the AI would know. to deliver that and if you desired a. sonnet about the indictment of Donald. Trump well it could Dash that off. instead. in late June I began to experiment with. a plug-in the Wikimedia Foundation had. built for chat GPT. at the time this software tool was being. tested by several dozen Wikipedia.
editors and Foundation staff members but. it became available in mid-july on the. open AI website for subscribers who want. augmented answers to their chat gbt. queries. the effect is similar to the retrieval. process that Jesse Dodge surmises might. be required to produce accurate answers. GPT 4's knowledge base is currently. limited to data it ingested by the end. of its training period in September 2021.
a Wikipedia plugin helps the bot access. information about events up to the. present day. at least in theory the tool lines of. code that direct a search for Wikipedia. articles that answer a chatbot query. gives users an improved combinatory. experience the fluency and linguistic. capabilities of an AI chatbot merged. with the factuality and currency of. Wikipedia. one afternoon Chris Alban who's in.
charge of machine learning at the. Wikimedia Foundation took me through a. quick training session. Alban asked chat gbt about the Titan. submersible operated by the company. Ocean Gate whose whereabouts during an. attempt to visit the Titanic's wreckage. were still unknown. normally you get some response that's. like my information cut off is from 2021. Alban told me but in this case chat GPT.
recognizing that it couldn't answer. albin's question what happened with. ocean Gate's submersible directed The. plug-in to search Wikipedia and only. Wikipedia for text relating to the. question. after the plugin found the relevant. Wikipedia articles it sent them to the. bot which in turn read and summarized. them then spit out its answer. as the responses came back hindered by. only a slight delay it was clear that.
using the plugin always forced chat GPT. to append a note with links to Wikipedia. entries saying that its information was. derived from Wikipedia which was made by. volunteers. and this as a large language model I may. not have summarized Wikipedia accurately. but the summary about the submersible. struck me as readable well-supported and. current a big improvement from a chat. GPT response that either mangled the.
facts or lacked real-time access to the. internet. Albin told me it's a way for us to sort. of experiment with the idea of what does. it look like for Wikipedia to exist. outside of the realm of the website so. you could actually engage in Wikipedia. without actually being on wikipedia.com. going forward he said his sense was that. the plug-in would continue to be. available as it is now to users who want.
to activate it but that eventually. there's a certain set of plugins that. are just always on. in other words his Hope was that any. chat GPT query might automatically. result in the chat Bots checking facts. with Wikipedia and citing helpful. articles. such a process would probably block many. hallucinations as well. for instance because chat Bots can be. deceived by how a question is worded.
false premises sometimes elicit false. answers or as Albin put it if you were. to ask during the first lunar Landing. who were the five people who landed on. the moon the chatbot wants to give you. five names. only two people landed on the moon in. 1969 however Wikipedia would help by. offering the two names Buzz Aldrin and. Neil Armstrong and in the event the. chatbot remained conflicted it could say.
it didn't know the answer and linked to. the article. the plugin still lets chat gbt get. creative but in limited ways. the following week when I asked that for. updates about the Ocean Gate submersible. I got a three paragraph rundown of how. the tragedy unfolded including the. deaths of five passengers then I asked. it to formulate its answer and five. bullet points which it did instantly. could it then adapt those five bullet.
points I asked so that a seven or eight. year old could understand. here's a simpler version chat gbt said. instantly and offered just what I asked. for noting that the Titan was a special. underwater vehicle and its implosion was. a sad event. it wasn't perfect I told Chachi BT that. its bullet points seem to overlook how. Stockton Rush ocean Gate's chief. executive had been criticized for. ignoring safety standards.
you raise a valid point it responded. here's a revised version that addresses. your concern. its fix took only a few seconds. within the Wikipedia Community there is. a cautious sense of hope that AI if. managed right will help the organization. improve rather than crash. Selena deckleman the chief Tech officer. expresses that perspective most. optimistically. what we've proven over 22 years now is.
we have a volunteer model that is. sustainable she told me I would say. there are some threats to it is it an. insurmountable threat I don't think so. the longtime Wikipedia editor who wrote. Death of Wikipedia told me that he feels. there is a case to be made for a good. outcome in the coming years even if the. longer term seems far less certain. the Wikimedia plugin is the first. significant move toward protecting its.
future. projects are also in the works to use. recent advances in AI internally. Albin says that he and his colleagues. are in the process of adapting AI models. that are off the shelf essentially. models that have been made available by. researchers for anyone to freely. customize so that Wikipedia's editors. can use them for their work. one focus is to have ai models Aid new.
volunteers say with step-by-step chat. bot instructions as they begin working. on new articles a process that involves. many rules and protocols and often. alienates Wikipedia's newcomers. Laila Zia the head of research at the. Wikimedia Foundation told me that her. team was likewise working on tools that. could help the Encyclopedia by. predicting for example whether a new. article or edit would be overruled or.
she said perhaps a contributor doesn't. know how to use citations in that case. another tool would indicate that. I asked whether it could help Wikipedia. entries maintain a neutral point of view. as they were writing absolutely she says. for the moment as the Wikipedia. Community debates rules and policy. article submissions entirely written by. llms are heavily discouraged on English. language Wikipedia.
still there remains a kind of John Henry. problem with AI. the chat Bots unlike their human. counterparts have a formidable ability. to churn out language like a steam. driven machine 24 7. I suspect the internet is going to be. filled with crud just all over the place. Chris Alban told me. and with the AI models getting better at. mimicking people's writing styles it may. be increasingly difficult to detect chat.
bot written submissions. one Wikipedia editor whose first name is. Theo sent me links in early June to show. how he was in the midst of fending off a. barrage of edits involving suspect. citations formulated by AI including one. to an article about Lake boksa in Greece. often I got the sense that Theo and. other wikipedians were worried that. their human abilities to scrutinize new. content and citations stretched to the.
Limit already might soon be overwhelmed. by an avalanche of AI generated text. certainly new tools that were themselves. AI would help but even if the editors. won in the short term you had to wonder. wouldn't the machines win in the end. three years ago in anticipation of. Wikipedia's 20th anniversary Joseph. Regal a professor at Northeastern.
University wrote a historical essay. exploring how the death of the site had. been predicted again and again. Wikipedia has nevertheless found ways to. adapt and endure Regal told me that the. recent debates over AI recall for him. the early days of Wikipedia when its. quality was unflatteringly compared to. that of other encyclopedias. it served as a proxy in this larger. culture War about information and.
knowledge and quality and authority and. legitimacy. so I take a sort of similar model to. thinking about chat GPT which is going. to improve just like Wikipedia is not. perfect it's not perfect it's never. going to be perfect but what is the. relative value given the other. information that's out there. the future as he saw it would be a range. of options for information caveat mtor. including everything from chat gbt to.
Wikipedia to Reddit to tick tock a. dedicated plugin could meanwhile improve. the chatbot's answers to questions about. for instance Health weather or history. at the moment it goes against the grain. to bet against AI the big tech companies. wagering billions on the new. technologies and largely undaunted by. their shortcomings or risks seem intent. on forging ahead as fast as they can.
those Dynamics would suggest that. organizations like Wikipedia will be. forced to adapt to the Future that AI. has begun to create rather than exert. influence over AI or mount an effective. resistance to it. yet many wikipedians and academics I. spoke with question any such assumption. impressive as the chat Bots may be ai's. apparent Glide path to success May soon. encounter a number of obstacles.
these could be societal as well as. technical the European Union's. Parliament is presently considering a. new regulatory framework that among. other things would Force tech companies. to label AI generated content and to. disclose more information about their AI. training data. Congress is meanwhile considering. several bills to regulate AI. legal scrutiny may be coming to in one.
closely watched lawsuit stability AI is. being challenged for using pictures from. Getty Images without permission. a California class action suit accuses. open AI of stealing the personal data of. millions of people that has been scraped. from the internet. while Wikipedia's licensing policy lets. anyone tap its knowledge and text to. reuse and remix it however they might. like it does have several conditions.
these include the requirements that. users must share alike meaning any. information they do something with must. subsequently be made readily available. and that users must give credit and. attribution to Wikipedia contributors. mixing Wikipedia's Corpus into a chatbot. model that gives answers to queries. without explaining the sourcing May thus. violate Wikipedia's terms of use two. people in the open source software. Community told me.
it is now a topic of conversation inside. the Wikimedia Community whether some. legal recourse exists. data providers may be able to exert. other kinds of Leverage as well. in April Reddit announced that it would. not make its Corpus available for. scraping by big tech companies without. compensation it seems very unlikely that. the Wikimedia Foundation could issue the. same dictum and close its sites off an. action that Nicholas Vincent has called. a data strike because it's terms of.
service are more open. but the foundation could make arguments. in the name of fairness and appeal to. firms to pay for its API just as Google. does now. it could further insist that chat Bots. give Wikipedia prominent attribution and. offer citations in their answers. something Selena deckleman told me the. foundation is discussing with various. firms. Vincent says that AI companies would be.
foolhardy to try to build a global. encyclopedia themselves with individual. contractors instead he told me there. might be an intermediary stage here. where Wikipedia says hey look at how. important we've been to you. such an entreaty Could Be an Effective. reminder too that the chat Bots are made. from us. without ingesting the growing millions. of Wikipedia Pages or vacuuming up. Reddit arguments about plot twists in. the bear new llms can't be adequately.
trained. in fact no one I spoke with in the Tech. Community seemed to know if it would. even be possible to build a good AI. model without Wikipedia. it may require the equivalent of a death. in the family before the tech companies. realize that they exist in a world of. mutual dependency. already according to the computer. scientists working in the AI industry. some technologists are concerned that.
new AIS are compromising the health of a. website for programmers called stack. Overflow a popular platform that the. models have been trained on to answer. coding questions. the problem seems to have two distinct. aspects if those with coding inquiries. can go to chat GPT for help why go to. stack overflow. in the meantime if fewer people are. Consulting stack Overflow for answers. why continue posting helpful suggestions.
or insights there. even if conflicts like this don't impede. the advance of AI it might be stymied in. other ways. at the end of May several AI researchers. collaborated on a paper that examined. whether new AI systems could be. developed from knowledge generated by. existing AI models rather than by human. generated databases. they discovered a systemic breakdown a. failure they called Model collapse.
the authors saw that using data from an. AI to train new versions of AIS leads to. chaos. synthetic data they wrote ends up. polluting the training set of the next. generation of models being trained on. polluted data they then misperceive. reality. the lesson here is that it will prove. challenging to build new models from old. models. and with chat Bots Ilya shumailov and.
Oxford University researcher and the. paper's primary author told me the. downward spiral looks similar without. human data to train on shimailov said. your language model starts being. completely oblivious to what you ask it. to solve and it starts just talking in. circles about whatever it wants as if it. went into this madman mode. wouldn't a plug-in from say Wikipedia. avert that problem I asked it could she.
my love said but if in the future. Wikipedia were to become clogged with. articles generated by AI the same cycle. essentially the computer feeding on. content it created itself would be. perpetuated. ultimately the study concluded that the. value of data from genuine human. interactions will be increasingly. valuable for future llms. at least for today's wikipedians that. seems like encouraging news insofar as.
it suggests our new machines will need. us at least for a while to keep them. honest and functional and dependent on. us. ensuring that an AI system is doing. what's in the best interests of humanity. involves a theoretical concept known as. alignment. alignment is viewed as both an enormous. Challenge and an enormous priority for. AI because a system out of sync with.
humans might create terrible damage. if AI ruins or compromises a mostly. reliable system of free knowledge it's. difficult to see how that aligns with. our best interests. one of the things that's really nice. about having humans do the summarization. is that you get some sort of basic level. of alignment by default Aaron hafaker. pointed out to me. and if you appreciate the editors of. Wikipedia are human they have human.
motivations and concerns and that their. motivations are providing high quality. educational material to align with your. needs then you can essentially put trust. in the system. you can grasp the alignment argument. better when you talk to people who. devote their lives to the idea. when I asked Jade who has more than 24. 000 edits to her credit why she spends. her free time typically 10 to 20 hours a. week editing Wikipedia she said she.
believed in Sharing knowledge. plus I'm just a big nerd she said we. were speaking by Zoom late in the. evening and it was a conversation that. had little resemblance to other long. evenings of dialogue I'd had with chat. GPT. some of Jade's work spoke to her. personal interests in nature and birds. like an entry she wrote on the. Vermillion flycatcher which got about 21. 000 page views in the past 12 months.
she also told me she works regularly on. the Wikipedia entry on the American. Civil War which had 4.84 million views. over the same period. her goal was to continue to work toward. completeness and greater accuracy in. that Civil War article so that it. achieves featured status on Wikipedia a. rare recognition usually marked by a. star of an article's quality that is. awarded by Wikipedia's editors to about. 0.1 percent of English language entries.
my calculations in the past are you know. more than 10 million people read my work. in a year Jade said so it's an honor to. have people reading all that we are. going to have to create processes we are. going to have to have hard conversations. she said about the ethics of using AI to. create Wikipedia articles. when I asked her whether chatbots would. soon eliminate her opportunities for.
volunteer work she replied I don't ever. maybe none never but certainly not in. this Century do I see robots fully. replacing humans on Wikipedia I wasn't. as sure. the Allure of a chatbot conversation. despite its factual shortcomings already. seemed too irresistible and too. enchanting to too many millions of. people. in fact my own hours spent with chat GPT.
had chipped away at my own Mutual point. of view. not because the informational exchange. was so rigorous and detailed it wasn't. but because the interaction was so. captivating and effortless. nevertheless Jade was resolute. I'm an optimist she said. [Music]. thank you.
