By saying that a human edited and approved every part of the LLM generated content, changes things quite a bit. Yes, a person in 1995 would most likely enjoy the writing because it was ultimately reviewed and approved by human sensibilities. Even today, a reader would be hard-pressed to identify AI-generated content if it was properly edited by a conscientious author.
The problem isn’t that an AI/LLM was used (though some people oppose it in any capacity), it’s that some authors believe it can produce final content without the need of human oversight. As soon as you remove human oversight, it’s currently impossible to guarantee that the content serves the game/story in a meaningful, purposeful way. You seem to agree with this sentiment with how you clarified your thought experiment. So I guess we agree?
I assume where we disagree is the tone of how that sentiment is conveyed. I’m very hypercritical of LLM usage. You’re not as critical as me, it would seem. But I think we would agree on more things about LLMs than either of us may realize.
I read read a bigger moral question into this. It reads to me as saying, if something can pass the turing test, is it ok to present it to someone without telling them it’s artificial? I would argue that it’s not, though that’s a moral statement.
And I see creativity as an important cog in society. You can say artistic intent might be pretentious, but I see intent as communicating an idea. Once the communication of our ideas becomes auto-generated, we stop thinking critically about them. Our ideas basically stop at “It’d be cool if…” and then we live out the plot of Idiocracy.
This is a really brilliant observation. I have often but idly thought about the usefulness of an MCP layer that would take a player’s input and distill it into what is possible in a parser game. But you are right. It is very unlikely that a player is looking to type more stuff in, especially after the novelty wore off and they realized they could do things like “x door” or “look knife.”
Candy64: Agreed. Using it as a crutch is no good. Ultimately it’s replacing nothing although I doubt the toothpaste is going back into the tube, so everyone is going to have to make some kind of peace with it.
Daniel Stelzer: I agree with all of this. The kind of prose I was referring to would be human approved as opposed to human authored. Which may still not be enough, but I would never call it “AI slop”. True AI slop is a serious problem that should be fought. Whether human-approved prose should be allowed in Comps is debatable (and I wouldn’t complain if it’s banned). I have the utmost respect for human-authored works (and code).
HAL9000: Agreed with everything you said. We are not far apart at all.
Ivan Cockrum: Good point. That’s something for me to ponder. I can imagine differing opinions here.
I am very pragmatic, perhaps overly so with this technology. But I also think it is likely to be absolute garbage, at least western LLM tech. Rumored Chinese LLMs running with equal compute at 10% the power, if true, confirm it along with absurd data center buildout. We are apparently being heavily subsidized when using it. Ultimately, the entire argument may be moot when it implodes, as they must demand large fees to access it. Hobbyist usage (including in IF) goes out the window; I know I won’t pay a nickel for it. Cheap, super-efficient local LLMs seem like the future. The problem may solve itself if very few IF developers are running them.
A new, unknown author publishes a promising first novel that’s generally well-received. Because nobody has noticed that it’s just Henry Green’s 1939 novel Party Going with the names changed.
Someone is cleaning their attic and discovers the manuscript of an unpublished Dickens novel. It’s published to great acclaim, and nobody even notices that the house it was “discovered” in doesn’t even have an attic.
Next door, the owner of that house “discovers” the dramatic personal diary of a Union soldier killed at Antietam, which reveals a wealth of “knowledge” about the day-to-day life of a soldier of the day.
Next door to that, the “discovery” is a diary contemporaneous with the one above, but gives the perspective of an enslaved person who escapes captivity via the underground railroad, providing details and insight never previously known.
Same example as above, but in an alternate universe where the “discoverer” is an ardent segregationist, and points to the “diary’s” relatively positive depiction of the institution of chattel slavery as a defense of their political views.
A couple doors down, what is “discovered” isn’t a novel or a diary, but a 2000 year old, previously unknown gospel written by one of the apostles.
Presumably at least some of those examples don’t provoke the reaction of “well, if nobody reading it can tell, why does it matter?”
I also think that, entirely independent of the whole “attribution matters, even if you couldn’t tell” argument, a lot of people here probably are exceedingly skeptical of the framing insisting on the complete indistinguishability of genAI prose. Because I think literally everyone who has been online for the past year or so at some point has had the experience of having to scroll past a bunch of absolutely godawful genAI slop. And quite a few have probably sampled some number of genAI games, which are, characteristically, dire.
So I think it is the case that, no, people really do care about where stuff comes from andseparately are probably rightly skeptical of the assumptions the proposed conundrum makes. And I say this as someone who isn’t ideologically opposed to genAI.
Your points are well made. If I understand correctly, it sounds like those scenarios involve dubious practices like subterfuge or rewriting history, which I never endorsed. I sidestepped the ethical considerations to simply focus on whether Gen-AI prose could be suitable for an IF work of passable quality. I agree that using Gen-AI should carry disclosure for people who are adamant they never want to deal with it - even if it falls into the bucket of “human approved” (which I argued legitimately exists and is not automatically AI slop). True AI slop has caused endless pain and is indefensible.
I understand that’s what you’re trying to do. But that presumes that the ethical considerations are separable, which I’m not sure they are. And it’s clear that many people tacitly believe that they are not.
And I also don’t think it’s necessary to assign malicious intent for there to be a distinction made. If someone simply mistook a painting by some random and very unfamous artist for a van Gogh (or vice versa) they’re going to adjust their perception of the value of the work in accordance with that belief regardless of whether or not theyd been steered into it through deliberate deception. That is, the locus of the shift in perceived value is in the belief in the provenance qua the provenance itself. I think they would be additionally outraged if they discovered they had valued something because someone else had mislead them, but that’s a separate thing. As near as I can tell, your argument is basically just that because the two things might not be distinguishable to someone, therefore their value is similar. And my point is that people absolutely do not appear to feel that way in general about things.
Say you’re a baseball player and you win the Big Game by hitting a home run. You’re not famous, it’s not even the major leagues. So it’s far more important to you than anybody else. You’re given the game-winning ball as a souvenir. One day when you’re cleaning your house you accidentally drop it into a box of more or less identical baseballs. Nobody has deliberately mislead you. One of the balls is certainly the one you hit the homer with. Do you just pick a random one, and designate it as the trophy to keep on your desk, because you’re not fooling anyone or anything, it’s just a memento there to remind you of your accomplishment? Do you keep all, let’s say, ten balls from the box, put them all in display cases, assigning 10% of the nostalgia to each? Just write off the whole thing, figuring that even though one of those ball is absolutely the “special” one, there’s just no way of determining which it is, so they’ve all become “worthless” as a memento?
I think whatever you decide about the baseball, it’s entirely going to be based on some vague intangible something whose locus is more or less entirely in your perceptions of it. It’s not that one has more essential “baseball-ness” or is better suited for play or anything like that. And I don’t think that’s a weird special case or contrived example or anything like that. It’s just a basic fact of how people value things.
In other words, people place different values on things they couldn’t tell apart in a blind test all the time. Because people aren’t perfectly logical machines that are just trying to solve min/max problems to maximize their utility or something like that. A bewildering amount of human activity is devoted to pursuits denominated in intangibles.
I feel like something missing from this argument is reproduction. I probably didn’t notice the first AI prose I read and maybe not the second.
My “AI slop detector” didn’t kick in until I started to see the same patterns reproduced over and over in different contexts.
So taking one piece of LLM authored prose back to 1995 it might seem indistinguishable. But if you took all the LLM authored prose back then I think people might not understand why it was happening but they would notice when it was all over the place.
Perhaps the AI issue is not about prose quality. If prose is bad, it’s bad - who cares if it’s “natural”? And AI doesn’t always mean bad. If folks employ AI to improve their own prose, I’m all for it.
As long as the human author acts as the chief game designer, and the result looks great, I have no issue with AI whatsoever.
When we are in the next AI era where chatbots can actually design good games all by themselves… well, that’s another story.
But as long as you are the author, and it’s your vision being brought to life by any means available, you’re fine.
Around 22 years ago there was a crummy adaption of Asimov’s I, Robot to the silver screen. It had a couple of bits of good dialogue though:
Detective: Human beings have dreams. Even dogs have dreams, but not you, you are just a machine. An imitation of life. Can a robot write a symphony? Can a robot turn a… canvas into a beautiful masterpiece? Robot: Can you?
I’d argue (as others here have) that while humans may also indeed produce terrible prose, it is still produced by a human. E.B. White famously said writing is an act of faith, not a trick of grammar. You reveal the person you are by what you write.
At least 4 of your examples wouldn’t pass Authentication or Catalogue Raisonné tests, so they wouldn’t likely get to the stage where a member of the general public would have had to consider how they would react.
And in these modern time many members of the general public can’t tell real from AI generated anyway, as the 2023 AI vs Human: The Creativity Experiment documentary demonstrated. And generative AI has advanced, and some would say people in general have potential digressed, in the 3 years since that documentary was released.
It’s not all about brevity, though, it’s about vocabulary. I find parser games almost unplayable, and I’ve come to realise that a good part of that is I just don’t use the verbs the designer is expecting, I use a wide variety of synonyms instead. If there was something capable of recognising all those synonyms — without the author having to think of them all first — that might be nice. (Though I also realise I loathe parser-style puzzles, so maybe not).
I’m so used to the characteristic rhythms and repetitive phrasing of the AI bots Quora now uses to answer 50% of their questions (as well as write them) that I don’t even have to glance up at their Noun+Noun user names to confirm.
This is one unexplored direction for using LLMs in parser games: the game writer using them to generate every conceivable English synonym and alternative expression for every command, and shipping all of them with the game (likely as a multi-gigabyte game file, mind you.)
Then the need for the player to wait for a slow live connection is eliminated, for one thing. Has anyone tried anything along these lines?
Even reading this very narrowly, as an argument about the prophylactic effect of the catalogue raisonné approach, I disagree with your confidence in its accuracy. I also do not share your implicit confidence in the universal application of such standards before publication.
But I think that’s really all secondary, because just asserting “that’s unlikely to happen” isn’t really a responsive counterargument, as the argument isn’t about, or contingent upon, a particular estimation of the likelihood of any of the proposed scenarios.
I’m positively scandalized that members of the general public are being exposed to works of art that haven’t undergone a rigorous catalogue raisonné evaluation!
This was already done in the old rules-based AI systems from last decade. You write a rule that checks for ~get and it’s checking for every synonym of ‘get’ in this pre-built library. It might not have every single synonym, but it’s pretty good.
You would have to tie those together manually, but it gives you the framework. Thanatophobia, which was mentioned above, is the only example I know of a game that used this system. In this case, it was a purely dialogue driven game. But there’s no reason why you can’t make an adventuring parser out of it. For your example you could say: