I, Robot, Am a Camera.
If you read the news, or watch the news, or have been in the presence of the internet at any point in the last few weeks, you have probably heard that AI is planning to kill us all.
It's a rather outlandish prediction, and one that is notably short on particulars or specifics. Indeed, a few news outlets are finally beginning to ask the obvious questions - like how exactly is a Large Language Model supposed to exterminate eight billion human beings...? But these questions are muted and cautious after nearly two weeks of largely uncritical reporting.... and much hysteria on the part of the general public.
"As people warn that A.I. could destroy humanity, they are light on examples of how," said the New York Times over the weekend. And the BBC quoted an unnamed tech worker who said that their "amusement at the new moment of existential AI fears largely stemmed from how little detail had been provided by its proponents to defend the notion that all of human life was at stake. The claims are 'always vague', the person said."
Vague or not, these apocalyptic claims have been embraced by a very receptive public (at least in the West) partly because they are just the latest in a long line of unremittingly hostile stories about the harm wrought by ChatBots. We have already been told that AI has led to "social de-skilling". We have been told that AI data centres are using up our water (although lots of other industries use a lot more water without nearly as much outrage).
And most recently (throughout much of August) the news cycles have been dominated by horror stories about AI companies buying up massive stockpiles of "rare books" to be scanned and then destroyed in a never-ending quest to "train" ChatBots on the stolen labours of our greatest writers and thinkers.
This particular story had been dominating the airwaves for most of the Summer, and has only recently been drowned out by the more recent hysterical warnings about AI Apocalypse.
This is one of the difficulties about modern AI; the news stories are shifting so fast that it's difficult to talk about anything before the focus moves on to some new techno-atrocity. In just a few weeks we've gone from "AI is Burning Books" to "AI is Planning to Kill Us All."
We've been down this road before with other apocalypse-level villains. If the pattern holds, we can expect to see "AI Has Invaded Poland" sometime around Halloween.
| In a rare moment of accord; political cartoonists of every stripe are happy to board the AI-apocalypse train. |
It's a shame in a way, because unlike the Killer-Robots-Are-Coming-For-You screeds, the stories about book-scanning are actually true, and are unambiguously happening right now. AI companies really are purchasing vast quantities of old books on the second-hand market. They really are scanning them, processing their contents as "training data" for their Large Language Models, and then destroying them (the books, I mean; not the Large Language Models). It's all true.
Except it isn't as bad as you think.
I'm going to ask for your patience while I delve into this subject, because it touches on some important concerns -and some equally important misconceptions - about the nature of AI.
To begin with, we need to talk about a really bad adaptation of The Time Machine, by H.G. Wells.
Part the Second: What a Movie...
The 2002 adaptation of The Time Machine is, in virtually every respect, a terrible, awful movie. The plot is mawkish and incoherent, the dialogue clunky and awkward. The production design is bland and unimaginative - especially when compared with the joyfully bonkers 1960 version. And there are several moments that stretch the limits of credulity, even for a Hollywood action flick.
| To top it all off, the moon crashes into the Earth, a little bit. At least a volcano doesn't erupt... |
If you ever have an occasion to watch this film, just... don't. Seriously; it's a crime what they put on the screen. I could hardly believe what I'd seen.
All of its escapist Technicolor twaddle notwithstanding, the 2002 Time Machine did feature a curious little scene in New York-of-the-near-future (about 2030, we are told). The time traveller, exploring this strange future city, finds his way into the New York Public Library, where he encounters a walking, talking, singing (yes, really) holographic AI interface of the entire library database.
The AI interface (played in the film by Orlando Jones) proudly boasts that "he" is a "compendium of all human knowledge"; something he explicitly (and very implausibly) demonstrates later in the film (and eight hundred thousand years in the future, plotwise).
Every book, every poem, every web post has apparently been scanned into his memory, to be retrieved at any moment for the benefit of anyone who should ask him. Very useful should you ever need to rebuild human society from scratch.
| This thing still works after eight hundred thousand years? I'm lucky if my phone will switch on after eighteen months... |
(Don't worry; I am not screening this film on Thursday. I will never screen this film on Thursday. Stay with me...)
When The Time Machine was released in 2002, no one had ever heard of a Large Language Model. The notion of a conversational ChatBot built out of the assembled totality of human language was, at that time, purely the stuff of science fiction. Now, over twenty years later, we are experiencing something approaching the reality of this idea, and many authors are extremely unhappy about it.
A number of lawsuits are currently working their way through the courts, brought by various authors and publishers, objecting to what they consider to be the unauthorised and unlawful use of their copyrighted material for the purpose of "training" these ChatBots.
As the author Charles Graeber recently put it in an opinion piece for the New York Times,
"[We] brought a class-action lawsuit against Anthropic. Our complaint was piracy and copyright violation; Anthropic had stolen our books because it needed them to create a commercial product, a generative A.I. chatbot designed specifically to write like us. But they didn’t ask us or pay us. That seemed unfair, and illegal.
Piracy is a crime as old as gold. […] And intellectual property theft is pretty much baked into the A.I. development story. Most, if not all, A.I. companies built their tech on piracy — though they may call it fair use."
Whether these lawsuits will have any success is an open question; right now it's still at the lawyers-talking-to-other-lawyers stage (no matter what transpires, these things are always great for lawyers) but the argument seems to be boiling down to the difference between reading a book and stealing it.
Part the Third: Humans Read Books
Let's assume for a moment that I like to read books. (Feel free to assume this. I do like to read books.)
When I learn about a book that I'm interested in, several things generally happen:
1. I try to find a copy of the book.
2. I obtain a copy of the book (by ordering it online; by purchasing it from a bookseller, or even by borrowing it from a library)
3. I read the book. (I hope you're all taking notes, by the way; there may be questions after)
4. I have read the book.
Step 4 is more important than you might think because once I've read a book, that book becomes part of me; part of my lived experience. That experience will (hopefully) transform me into something slightly more than I was before.
If it's a good book (and even if it isn't) it will intermix with all the other books I have read; all the movies I have seen; all the people I have interacted with throughout my (hopefully ongoing) life. The cumulative result of that unique combination of experiences is part of what makes me the person I am. When I read (for example) The Scarlet Letter, I am not experiencing Nathanial Hawthorne's words in a vacuum. Those words are taking their place in my consciousness alongside Rebecca and Billy Budd and Do Androids Dream of Electric Sheep and The Haunting of Hill House and absolutely everything else I have ever read in my life.
We are, each of us, the sum of our experiences, and every book we read makes us that little bit more experienced than we were before. And this of course is exactly why writers publish their works: they want their words to have an effect on their readers. "Your book changed my life" is the best thing you can ever say to a writer.
But here's the thing about reading: once an author's words have entered my system, they become part of my identity from that moment forward. I could sell my copy of the book; I could give it to a friend; I could leave it on the bus for someone else to find, or I could even throw it away if I really want to. The words will still be part of me, and the experience of having read them will be a part of what it means to be "me" from that point forward. If you want to get poetic about it, you could say that I have absorbed the "soul" of the book. It has become part of my essence.
This is what Philip K. Dick was talking about when he said that fiction "ultimately winds up being a collaboration between author and reader, in which both create - and enjoy doing it." (And he was speaking as a reader, not a writer when he made that observation)
But a "collaboration" involves two people, and the author cannot control what happens to a text once it enters someone else's head.
To publish a book is, to a certain extent, to relinquish control over it. When a book has been "released into the wild" so to speak, it can be received by anyone who chooses to do so. A book may be read by a university professor (and taught in a classroom). It might be read by an abused housewife, a corrupt politician or a convicted murderer. Each individual is, in a sense, reading a different book, because they are each reading it as the individuals they are at that moment.
And then of course the book might be read by someone who is inspired to write books of their own. This isn't plagiarism; it isn't theft and it isn't piracy: it's an author taking up their position in the larger literary tradition. Stephen King is a phenomenally successful novelist who has openly cited another novelist (Shirley Jackson) as his primal inspiration. Ray Bradbury has enthusiastically said the same things about Leigh Brackett. Every writer is shaped by their experiences, and those experiences include all the other books they have read over the course of their lives. Margaret Atwood exists in a world that includes Geoffrey Chaucer and William Shakespeare and the Brontë Sisters. Louise O'Neill exists in a world that includes Margaret Atwood and Stephen King and Ray Bradbury.
| Literature is a continuum. Every writer carries the legacy of the writers who brought them to this point. |
The point of all this is that once a book exists in the wider world (once it has been published) the author's control over the fate of that book is not infinite. The book is accessible to anyone who has the ability to read, and the author has limited powers to curtail that accessibility. An author can legally insist that their book not be translated into, let's say... Aramaic. But the author can't decree that their book is off limits to, for example, Presbyterians. Or haberdashers.
That's not a huge problem at the moment, because I am not currently aware of any authors who actually want to withhold their books from Presbyterians or haberdashers. But there are many, many authors who strenuously object to their books being read by computers.
This of course is where things get complicated. If publication is a "Licence to Read" does that license extend to non-humans? And does scanning a book into a giant computer system count as "reading"?
To even begin to answer those questions, we need to consider what actually happens when a piece of text is incorporated into a Large Language Model.
Part the Fourth: Robots Eat Books
| The logo of a facility in Las Vegas where books are scanned and "fed" into Large Language Models. Really. |
Large Language Models, we are often told, have been "trained" on vast amounts of human text. Basically, they are the accumulation of virtually everything humans have ever put into writing (give or take).
All of that raw data allows them to "predict" which words should follow which other words (by analysing the patterns in stupefyingly large quantities of human language) and before you know what's happening, your machine is talking to you.
This is where the (fictional) AI system in the 2002 Time Machine movie casts a long and unwelcome shadow. In the film, that AI "librarian" has every work of human literature at his fingertips, ready to be accessed with a simple request.
"Recite Huckleberry Finn," you tell it, and the avatar begins to spout Mark Twain.
"Now do The Da Vinci Code," you say, and it will regurgitate Dan Brown, but without the royalty payments.
But that's a movie, and it's about as realistic as Arnie-the-time-travelling-murder-bot in The Terminator. Large Language Models in real life do not actually have all that "training data" available to them in searchable form. Despite what some authors might think, a ChatBot is not like a giant filing cabinet with hundreds of thousands of folders marked "Agatha Christie", "Raymond Chandler", "Douglas Adams", "William Shakespeare" etc. all neatly labelled and collated in flagrant violation of authorship and copyright.
ChatGPT can't reach into its own internal structure to extract Captain Corelli's Mandolin any more than you or I could dip into our DNA to recite the bits that give us ten fingers, or a working pancreas.
I recently ran a small-scale (and very un-scientific) experiment with an assortment of Language Models. I presented them with carefully obscure passages from several different novels - the kinds of novels that would almost certainly have been included in their training data - and asked them to identify the works. They were unable to do so with any certainty, although they came up with some jolly sharp guesses in most cases. These things are not the animated piracy-engines that many authors imagine them to be.
A better way to visualise all that training data would be to imagine a vast ocean that has been gradually built up from water sourced from various diverse locations. One drop of water in that ocean is Shakespeare. Another drop is Chaim Potok. A third drop is Barbara Cartland, and so on and so on. Each individual drop might have been unique and identifiable when it went in, but there is no way to isolate a specific drop after it's been added to the entire ocean. Scooping a cupful of water from the ocean isn't going to give you a spoonful of J. G. Ballard; it's going to give you something unique that is a synthesis of everything; of the vast combination of voices that collectively formed the ocean in the first place.
This is where the act of reading differs for humans and machines. We humans live in a physical world of bodies and environments and objects (yes; and books). A Large Language Model has no direct access to any of that; it only has access to words. Those words - books; essays; blog posts; dirty limericks... all those vast, endless volumes of volumes - are the physical reality of a ChatBot. Collectively, you could call them the machine equivalent of a human's life experiences.
If I live in London (which I do) my experiences here shape the kind of person I am. Everything I do, say or think is (overtly or subtly) affected by the environment of London and my lived experiences in it. But that doesn't mean I have the ability to reconstruct London on demand, any more than a ChatBot could recite Tom Clancy at you. Scanning Patriot Games into a computer doesn't turn ChatGPT into a Tom Clancy-Bot.
This is important (which is why I'm taking my time with it, so thank you for staying with me) because the news has been awash recently with stories about AI companies buying up massive quantities of old books, scanning them and then destroying them as part of its training process.
"WHY IS ANTHROPIC DESTROYING BOOKS?" was the Guardian headline last month. "AI Companies Are Buying—And Destroying—Antique Books," said Forbes, and The Week was in agreement: "AI companies are destroying rare books." And a website called Hyperallergic didn't pull its punches: "AI Is the New Book Burning."
This all seems fairly horrifying, and using terms like "rare" and "antique" makes it sound as if these corporations are systematically shredding the Library of Alexandria so that ChatGPT can write your emails for you. This is not happening. (The Library of Alexandria was torched by Julius Caesar in 48 BC, by the way. OpenAI has a solid alibi for that whole afternoon.)
All of the Summer chatter about "rare books" and the "new book burning" frenzy ultimately can be traced back to a single article, published by a website called 404media. 404media investigated a series of reports from online second-hand booksellers who had been observing a curious spike in bulk orders of old books. They traced one such order back to a warehouse in Nevada where the books were "destructively scanned". (This is a common practice for large-scale book scanning: the spine of the book is sliced off so that the individual pages can be more easily read by the scanner. Digitising books in this way is neither new nor controversial. It just feels icky if you're a book-lover.)
The article speaks vaguely of "rare books" without going into details (which is why so many other websites have seized upon that term for their scare headlines) but later in the piece it explicitly says "AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs." [my emphasis added]
Avoiding books that don't have ISBN barcodes immediately and completely rules out any books published prior to 1970 (when the ISBN system was adopted) so we're not exactly talking about the inventory of 84 Charing Cross Road here (and anyway, 84 Charing Cross Road has been a McDonald's for almost twenty years now. The rare books are long gone...).
In fact, there is absolutely no need to buy and scan anything genuinely "antiquarian" because those books will all be out of copyright and freely available anyway. Dickens, Jane Austin, Thomas Hardy, Hermann Melville, right up to early Agatha Christie and Raymond Chandler are all now in public domain, and their works have been uploaded to websites like Project Gutenberg where they may be read by anyone and everyone - including any curious ChatBots.
Ironically, AI developers are also avoiding any books published in the last five years or so, because (wait for it...) there is too much danger that those books might have been written with AI assistance, and they don't want the ChatBots to start training on themselves. (There's a Monty Python sketch in there somewhere, but it should also come as good news to any current authors who don't want their words fed into the big, bad Moloch-Machines. Congratulations, young writers; your books are safe... thanks to AI!)
The books that are prime targets for AI "consumption" seem to be the mass-market publications of the late 20th and early 21st Century. And this is where we need to face up to an uncomfortable truth (one that has been glossed over by the recent "book-burning" stories): the resale market for these books is virtually non-existent, and has been for some time. Even charity shops are mostly refusing book donations these days, and many of them quietly and discreetly destroy the crateloads of unwanted books they can't afford to store. Prison libraries are also now rejecting most of the books they are offered.
| If you're a book-lover, you really don't want to know what happens to most of the books given to charity shops. |
You literally can't give books away any more.
Remember my "Step 4" of book reading in the previous section? That step leaves a lot of old books behind, and those books can really pile up when you consider society as a whole. There are at least four million copies of Captain Corelli's Mandolin floating around out there somewhere, along with fifteen million copies of Jurassic Park, eighty million copies of The Da Vinci Code and something like one hundred and twenty million copies of the first Harry Potter. If a single copy of each of those gets sliced open for scanning purposes (after being legitimately purchased on the second-hand market) that isn't exactly going to bring the literary world crashing down around us.
The "legitimately purchased" part is important, and it's the point that will surely feature very prominently as all of these copyright cases work their way through the courts. The books in question haven't been stolen, because they were bought and paid for with actual money - and many of these books will certainly be titles that the booksellers were all too happy to offload (there are only so many copies of the same John Grisham novel that a bookseller can shift). The books haven't been pirated, because the AI systems do not have them stored somewhere to be dispensed on demand (as I say, you're confusing ChatGPT with the guy from The Time Machine).
Like it or not, these books are being read, and the reader is growing from the experience. That's exactly what's supposed to happen when a book is published, and most authors are thrilled... when humans do it.
I've been talking quite a bit about The Time Machine in this (never-ending) piece. Well, The Time Machine was written by H.G. Wells in 1895, and is considered the first significant work of fiction to depict a vehicle that can travel through time. Virtually every subsequent time travel story ever written owes its existence to H.G. Wells, but Wells (who lived for another five decades) was never tempted to sue anyone.
| The 1960 adaptation of The Time Machine boasted what is still the most stylish time machine ever depicted in a film. The 2002 version can only dream of such a thing. |
You could argue that every tale of doomed, star-crossed lovers (of which there are many, many) carries the seed of Romeo and Juliet. Again, this isn't plagiarism, it's just writers who read. Shakespeare is part of our cultural water supply, along with H.G. Wells, Jacqueline Susann, Harold Robbins and every other writer who, for better or worse, contributed to what we collectively call our literary tradition.
If the "I" in AI really does mean "Intelligence" then we should just let the robots read the damn books. If we're going to start making summary decisions about which demographics are barred from reading our books, then, well, I guess we don't get to complain about "book burning", do we?
I know this has been a long and tortuous discussion, and I thank you for your patience; especially since I haven't even talked about Thursday's film screening as yet.
So now, if you'll pardon me, I think it's time for a musical number.
No, really.
Part the Finalth: Rashomon, in B-Flat.
You may recall that I screened Akira Kurosawa's Rashomon back in July. I said at the time that it demonstrated the power of language to create Reality. I also pointed out the massive influence it had on wider culture: its name became a psychological term, like the Oedipus Complex or Munchausen's Syndrome.
Well, it also gave us a Cole Porter musical.
Many of you will know that I always try to include at least one musical when I plan out these series, and Les Girls fits the bill beautifully. It tells the story of Barry Nichols (played by Gene Kelly) and his three co-stars who collectively form Barry Nichols and Les Girls, a cabaret act based in Paris.
It was Cole Porter's final film musical and Gene Kelly's last film as a contract player for MGM studios. The late 1950s was also the last dying gasp of the classic Hollywood "studio" era, and there are some telltale signs of a rapidly changing cultural climate. Several of the musical numbers are extremely risqué, even by Cole Porter standards; a testament to just how toothless the once-invincible Hays Office was becoming.
But quite apart from all that, Les Girls wouldn't exist had it not been for Rashomon, seven years earlier.
Just as Rashomon gave us multiple, highly contradictory accounts of the same events (as told by various witnesses) Les Girls adopts the identical format. Everything is shown in flashback, and we have no way of knowing which account (if any) has any bearing on the "true" events.
Like H.G. Wells and his time-travelling vehicle; like Shakespeare and his star-crossed lovers, Rashomon has entered the water supply. Kurosawa's film was seen by audiences who reacted to it; were moved by it and, in this instance, were inspired by it. This isn't plagiarism (no one in the history of ever has even considered Les Girls a rip-off of Rashomon) but it is a prime example of a work of art that creates ripples in the fabric of society: ripples that influence everything that comes after.
And anyway, I think we all deserve a Cole Porter musical. It's been that kind of Summer.
Enjoy!
Comments
Post a Comment