A.I.’s Original Sin
from New York Times I'm Michael babbaro. this is the. [Music]. daily. today a times investigation shows how as. the country's biggest technology. companies race to build powerful new. artificial intelligence systems they. bent and broke the rules from the. start my colleague Kate Mets on what he.
[Music]. [Applause]. uncovered it's Tuesday April. [Music]. 16th Kade when we think about all the. artificial intelligence products. released over the past couple of years.
including of course these chat Bots. we've talked a lot about on the show we. so frequently talk about their future. their future capabilities their. influence on society jobs our lives but. you recently decided to go back in time. to ai's past to its Origins to. understand the decisions that were made. basically at the birth of this. technology so why did you decide to do. that because if if you're thinking about.
the future of these. chatbots that is defined by their past. the thing you have to realize is that. these chatbots learn their skills by. analyzing enormous amounts of Digital. Data so what my colleagues and I wanted. to do with our investigation was really. focus on that effort to gather more data. we wanted to look at the type of data.
these companies were collecting how they. were gathering it and how they were. feeding it into their systems and when. you all undertake this line of reporting. what do you end up finding we found that. three major players in this race open AI. Google and. meta as they were locked into this. competition to develop better and better. artificial intelligence they were. willing to do almost anything to get.
their hands on this data including. ignoring and in some cases violating. corporate rules and waiting into a legal. gray area as they gathered this data. basically cutting Corners cutting. Corners left and right okay let's start. with open AI the flashiest player of all. the most interesting thing we found is. that in late. 2021 as open AI the startup in San.
Francisco that built chat GPT as they. were pulling together the fundamental. technology that would power that chatbot. they ran out of data essentially H they. had used just about all the respectable. English language text on the internet to. build this system and just let that sink. in for a bit I mean I'm trying to let. that sink in they basically like a.
Pac-Man on a old game just consumed. almost all the English words on the. internet which is kind of unfathomable. Wikipedia articles by the thousands news. articles Reddit threads digital books by. the millions we're talking about. hundreds of billions even trillions of. words wow so by the end of 2021 open aai. had had no more English language text.
that they could feed into these systems. but their Ambitions are such that they. wanted even. more so here we should remember that if. you're Gathering up all the English. language text on the internet a large. portion of that is going to be. copyrighted right so if you're one of. these companies Gathering data at that. scale you are absolutely Gathering.
copyrighted data as well which suggests. that from the very beginning these. companies a company like open AI with. chat GPT is starting to break bend the. rules yes they are determined to build. this technology thus they are willing to. venture into what is a legal gray area. so given that what does open AI do once. it as you had said runs out of English.
language words to mop up and feed into. this system so they get together and. they say all right so what are other. options here and they say well what. about all the audio and video on the. internet we could transcribe all the. audio and video turn it into text and. feed that into their systems interesting. so a small team at openai which included. its president and co-founder under Greg.
Brockman built a speech recognition. technology called whisper which could. transcribe audio files into text with. high accuracy H and then they gathered. up all sorts of audio files from across. the internet including audio books. podcasts oi and most importantly YouTube. videos of which there's a seemingly.
endless supply right fair to say maybe. tens of millions of videos according to. my reporting we're talking about at. least a million hours of YouTube. videos were scraped off of that video. sharing site fed into this speech. recognition system in order to produce. new text for training open ai's chatbot. and YouTube's terms of service do not. allow a company like open AI to do this.
YouTube which is owned by Google. explicitly says you are not allowed to. in Internet parlance scrape videos on. Mass from across YouTube and use those. videos to build a new application that. is exactly what open AI did according to. my reporting employees at the company. knew that it broke YouTube terms of.
service but they resolve to do it anyway. so Kate this makes me want to understand. what's going on over at Google which as. we have talked about in the past on the. show is itself thinking about and. developing its own artificial. intelligence model and product well as. open AI scrapes up all these YouTube. videos and starts to use them to build. their chatbot according to my reporting. some employees at Google at the very.
least are aware that this is happening. they are yes now when we went to the. company about this a Google spokesman. said it did not know that openai was. scraping YouTube content and said the. company takes legal action over this. kind of thing when there's a clear. reason to do so but according to my. reporting at least some Google employees. turned a blind eye to open AI activities. because Google was all also using.
YouTube content to train its AI wow so. if they raise a stink about what open AI. is doing they end up shining a spotlight. on themselves and they don't want to do. that I guess I want to understand what. Google's relationship is to YouTube. because of course Google owns YouTube so. what is it allowed or not allowed to do. when it comes to feeding YouTube data. into Google's AI models it's an. important distinction.
because Google owns. YouTube it defines what can be done with. that data and Google argues that it has. a right to that data that its terms of. service allow it to use that data. however because of that copyright issue. because the copyright to those videos. belong to you and I lawyers who I've. spoken to say people could take Google. to court.
and try to determine whether or not. those terms of service really allow. Google to do this there's another legal. gray area here where although Google. argues that it's okay Others May argue. it's not of course what makes this all. so interesting is you essentially have. one tech company Google keeping another. tech company open AI dirty little secret. about basically stealing from. YouTube because it doesn't want people.
to know that it too is taking from. YouTube and so these companies are. essentially enabling each other as they. simultaneously seem to be bending or. breaking the. rules what this shows is that there is. this belief and it has been there for. years Within These companies among their. researchers that they have a right to. this data because they're on a larger. mission to build a technology that they. believe will transform the world and if.
you really want to understand this. attitude you can look at our reporting. from inside meta and so what does meta. end up doing according to your reporting. well like Google and other. companies meta had to scramble to build. artificial intelligence that could. compete with open. AI Mark Zuckerberg is calling engineers. and Executives at all hours pushing them.
to acquire this data that is needed to. improve the. chatbot and at one point my colleagues. and I got hold of recordings of these. meta Executives and Engineers discussing. this problem how they could get their. hands on more data where they should try. to find it and they explored all sorts. of options they talked about license ing. books one by one at $10 a pop and.
feeding those into the model they even. discussed acquiring the book publishers. Simon and Schuster and feeding its. entire Library into their AI model But. ultimately they decided all that was. just too cumbersome too time. consuming and on the recordings of these. meetings you can hear. Executives talk about how they were. willing to run off shod over copyright.
law and ignore the legal concerns and go. ahead and scrape the internet and feed. this stuff into their models they. acknowledged that they might be sued. over this but they talked about how open. AI had done this before them that they. meta were just following what they saw. as a market precedent interesting so. they go from having conversations like. should we buy a publisher that has tons. of copyrighted material suggesting that. they're very conscious of the kind of.
legal terrain and what's right and. what's wrong and instead say nah let's. just follow the open AI model that. blueprint and just do what we want to do. do what we think we have a right to do. which is to kind of just gobble up all. this material across the Internet it's a. snapshot of that Silicon Valley attitude. that we talked about because they. believe they are building this transform. formative.
technology because they are in this. intensely competitive situation where. money and power is at stake they are. willing to go there but what that means. is that there is at the birth of this. technology a kind of original sin that. can't really be. erased it can't be erased and people are. beginning to notice. and they are beginning to sue these.
companies over. it these companies have to have this. copyrighted data to build their systems. it is fundamental to their creation if a. lawsuit bars them from using that. copyrighted data that could bring down. this technology.
we'll be right. back so Kade walk us through these. lawsuits that are being filed against. these AI companies based on the. decisions they made early on to use. technology as they did and the chances. that it could result in these companies. not being able to get the data they so. desperately say they need these suits. are coming from a wide range of places. they're coming from computer programmers.
who are concerned that their computer. programs have been fed into these. systems they're coming from book authors. who have seen their books being used. they're coming from publishing companies. they're coming from news. corporations like the New York Times. incidentally which has filed a lawsuit. against open Ai and Microsoft MH news. organizations that are concerned over. their news article being used to build.
these systems and here I think it's. important to say as a matter of. transparency Kade that your reporting is. separate from that lawsuit that lawsuit. was filed by the business side of the. New York Times by people who are not. involved in your reporting or in this. daily episode just to get that out of. the way exactly I'm assuming that you. have spoken too many lawyers about this. and I wonder if there's some insight.
that you can shed on the basic legal. terrain I mean do the companies seem to. have a strong case that they have a. right to this information or do. companies like the times who are suing. them seem to have a pretty strong case. that know that decision violates their. copyrighted materials like so many legal. questions this is incredibly complicated. it comes down to what's called fair use. which is a part of copyright law that.
determines whether companies can use. copyrighted data to build new things and. there are many factors that go into this. there are good arguments on the open AI. side there are good arguments on the New. York Times side copyright law says that. you can't take my work and reproduce it. and sell it to someone that's not. allowed but what's called fair use does.
allow companies and individuals to use. copyrighted Works in part they can take. Snippets of it they can take the. copyrighted works and transform it into. something. new that is what open Ai and others are. arguing they're doing but there are. other things to consider does that. transformative work compete with the. individuals and companies that suppli.
the data that own the copyrights. interesting and here the suit between. the New York Times company and open AI. is. illustrative if the New York Times. creates articles that are then used to. build a chatbot does that chatbot end up. competing with the New York Times do. people end up going to that chatbot for. their information rather than going to.
the times website and actually reading. the article that is one of the questions. that will end up deciding this case and. cases like it so what would it mean for. these AI companies for some or even all. of these lawsuits to. succeed well if these tech companies are. required to license the copyrighted data. that goes into their systems if they're.
required to pay for. it that becomes a problem for these. companies we're talking about Digital. Data the size of the entire. internet licensing all that copyrighted. data is not necessarily feasible we. quote The Venture Capital firm and dreon. harowitz in our story where one of their. lawyer says that it does not work for.
these companies to license that data. it's too expensive it's on too large a. scale H it would essentially make this. technology economically impractical. exactly so a jury or a judge or a law. ruling against open AI could. fundamentally change the way this. technology is built the extreme case is. these companies are no longer allowed to. use copyrighted material in building.
these chat Bots and that means they have. to start from scratch they have to. rebuild everything they've built so this. is something that not only imperils what. they have today it imperils what they. want to build in the future and. conversely what happens if the courts. rule in favor of these companies and say. you know what this is fair use you were. fine to have scraped this material and. to keep borrowing this material into the.
future free of charge well one. significant roadblock drops for these. companies and they can continue to. gather up all that extra data including. images and sounds and videos and build. increasingly powerful. systems but the thing is even if they. can access as much copyrighted material. as they want these companies May may. still run into a problem pretty soon.
they're going to run out of digital data. on the internet that human created data. they rely on is going to dry up they're. using up this data faster than humans. create it one research organization. estimates that by 2026 these companies. will run out of viable data on the. internet wow well in that case what. would these tech companies do I mean.
where they going to go if they've. already scraped YouTube if they've. already scraped podcast if they've. already gobbled up the internet and that. altogether is not. sufficient what many people inside these. companies will tell you including Sam. Alman the chief executive of open AI. they'll tell you that what they will. turn to is what's called synthetic data. and what is that that is data.
generated by an AI model that is then. used to build a better AI model it's AI. helping to build better AI that is the. vision ultimately they have for the. future that they won't need all this. human generated text they'll just have. the AI build a text that will feed. future versions of.
AI so they will feed the AI systems the. material that the AI systems themselves. create but is that really a workable. solid plan is that considered high. quality data is that good. enough if you do this on a large scale. you quickly run into problems as we all. know as we've discussed on this podast. these systems make mistakes they.
hallucinate they make stuff up they show. biases that they've learned from. internet data and if you start using the. data generated by the AI to build new AI. those mistakes start to reinforce. themselves right the systems start to. get trapped in these cue saacs where. they end up not getting better but. getting worse what you're really saying.
is these AI machines need the unique. Perfection of the human creative mind. well as it stands today that is. absolutely the case but these companies. have grand Visions for where this will. go and they feel and they're already. starting to experiment with this that if. you have an AI system that is. sufficiently. powerful if you make a copy of it if you. have have two of these AI models one can.
produce new data and the other one can. judge that data it can curate that data. as a human would it can provide the. human judgment so to speak so as one. model produces the data the other one. can judge it discard the bad data and. keep the good data and that's how they. ultimately see these systems creating. viable synthetic data but that has not. happened yet and it's unclear whether it.
will work it feels like the real lesson. of your investigation is that if you. have to allegedly steal data to feed. your AI model and make it economically. feasible then maybe you have a pretty. broken model and that if you need to. create fake data as a result which as. you just said kind of undermines ai's. goal of mimicking human thinking and. language then maybe you really have a.
broken model and so that makes me wonder. if the folks you talk to the companies. that we're focused on here ever ask. themselves the question could we do this. differently could we create an AI model. that just needs a lot less. data they have thought about other. models for. decades the thing to realize here is. that is much easier said than done we're. talking about creating systems that can. mimic the human. brain that is an.
incredibly ambitious task and after. struggling with that for decades these. companies have finally stumbled on. something that they feel works that is a. path to that incredibly ambitious goal. and they're going to continue to push in. that direction yes they're exploring. other options but those other options. aren't. working what works is more data and more.
data and more data and because they see. a path there they're going to continue. down that path and if there are. roadblocks there and they think they can. knock them down they're going to knock. them down but what if the tech companies. never get enough or make enough data to. get where they think they want to go. even as they're knocking down walls. along the way that does seem like a real. possibility if these companies can't get. their hands on more data then these. Technologies as they're built today stop.
improving we will see their limitations. we will see how difficult it really is. to build a system that can match let. alone surpass the human. brain these companies will be forced to. look for other options technically and. we will see the limitations of these. grandiose visions that they have for the.
future of artificial. [Music]. intelligence okay thank you very much we. appreciate it glad to be. [Music]. here we'll be right back. here's what else you need to know today.
Israeli leaders spent Monday debating. whether and how to retaliate against. Iran's missile and drone attack over the. weekend hery halevi Israel's military. Chief of Staff declared that the attack. will be responded to in Washington a. spokesman for the US state department. Matthew Miller reiterated American calls. for Restraint of course we to make clear. to everyone that we talk to that we want.
to see de deescalation that we don't. want to see a wider Regional War that's. something that's been but emphasized. that a final call about retaliation was. up to Israel Israel is a sovereign. country they have to make their own. decisions about how best to defend. themselves what we always try to and the. first criminal trial of a former US. president officially got underway on. Monday in a Manhattan courtroom Donald. Trump on trial for allegedly falsifying.
documents to cover up a sex scandal. involving a porn star watched as jury. selection began the initial pool of 96. jurors quickly dwindled more than half. of them were dismissed after indicating. that they did not believe that they. could be impartial the day ended without. a single juror being. chosen today's episode was produced by. Stella tan Michael Simon Johnson muj.
Zade and Ricky nety it was edited by. Mark George and Liz o Balin contains. original music by Diane Wong Dan Powell. and Pat mccusker and was engineered by. Chris Wood our theme music is by Jim. brunberg and Ben lek of. wonderly that's it for the daily I'm. Michael baboro. see you tomorrow.
[Music].
