Classifying and Labeling AI-Developed Works [Developed without AI]

I agree. I don’t think output from an LLM is “human made”. However, nobody is saying anything contrary to that. That isn’t the point of contention.

You have this fascination with the concept of any digital tools (outside of an assembler) somehow tainting the worthiness of the “human-made” label. Does this logic apply to all digital content, like art, music and animation? I mean, the image data is compiled by the art program; not to mention that those digital brushes can take a lot of load off the artist with pixel manipulation to render realistic brush strokes.

Is digital art not considered “human-made”? If so, do AI-generated images sit on the same level as not being human-made as images made in an art program? If not, then now you’re seeing the distinction that others are talking about when it comes to digital content. If you don’t see that distinction, I won’t pester you any further.

That is inaccurate.

For each output token, the LLM produces a probability distribution, which is basically a list of all possible tokens that could come next, with weights attached: “bite” with weight X, “daffodil” with weight Y, and so on.

But that doesn’t mean the system has to give you, the user, a sequence of tokens generated at random according to those probabilities. That’s where temperature comes in. At temperature zero, the system uses greedy decoding: it picks the highest-weighted, “most likely” token every time. Higher temperature settings increase the chance that you’ll get one of the less-likely tokens as output.

Notice that turning the temperature down to zero doesn’t necessarily make it work any less well. Mainly, it’s just boring. People don’t want to always get the same response every time when they’re talking to a chat bot. But for tasks where originality isn’t the point, it can be fine. OpenAI’s documentation recommends using temperature 0 for use cases like data extraction and factual Q&A.

That’s an interesting question.

In the case of image data being compiled by the art program, I’d say the file isn’t human-made, but the abstract picture it encodes is.

Similarly, at a high enough level of abstraction, the logic in a program is human-made, even if at a lower level it isn’t. Asking Claude to achieve a goal and letting it figure out which algorithm to use, asking Claude to implement a specific algorithm in C, writing the C code yourself, writing the assembly code yourself, and typing in the raw numeric opcodes all exist on a continuum – and at any point on that continuum, you’ll find people who believe their neighbors on one side are letting machines do all the work for them and their neighbors on the other side are stubbornly refusing to use perfectly good tools.

In the case of digital brush strokes, I’d say it comes down to whether the program is simply “taking a load off the artist” by automating some tedious pixel work that they could (and would) do themselves, or whether it’s creating art that otherwise couldn’t or wouldn’t be created through the artist’s own effort.

Less of the content of an AI-generated image is “human made” than of an image made in an art program. But there’s a good chance that neither of them qualify as fully “human made”, and I wouldn’t advise artists to put such a label on their work either. The more you apply things like filters and content-aware fill, the more of your artwork is being made by a machine rather than a person – not that there’s anything wrong with that.

As always, there’s a relevant XKCD: xkcd: Real Programmers

I somewhat agree, but I don’t think it’s that simple. Trying to sculpt the Venus de Milo out of marble using no tools would be quite the feat. Does that mean the original wasn’t human-made? Again I find myself coming back to the idea of leverage. A tool which enhances a creator’s ability through multiplication of effort such as speeding things up, applying more force, more precision, or “taking the load off the artist” in your words, doesn’t fundamentally alter the human nature of the creation; and yes I would put compilers in that category. I don’t think most people would think to label a 3d print as “human-made” even if they might label the original model as such and I agree with you that some art program features probably cross a line in the same way many physical-world automated tools do.

I’ve been trying to think about all this in the context of music sampling. It feels like the same thing, the same cultural crossroads, and I have the historical perspective of living through music sampling. When music sampling became accessible to the masses and inexpensive, there was great panic that sampling was going to destroy music as we knew it. That never happened. It profoundly altered music. It didn’t destroy it. It did take decades for this new technique of taking another artists work and reproducing it to finally get sorted out. I think this Deja-vu is happening again. What is AI/LLM usage? The ability to sample from other creators work and synthesize something new from it. That moment and this moment feel like the same societal crossroads. If history is any kind of teacher that she should be then we’ll all adapt and the future, when it comes for us all, and it may not be all doom and gloom. I’m not anti-AI but I am 100% in agreement with the cheese pizza problem.

I like the concept of labeling AI usage on a scale from 0-5 instead of all-or-nothing as outlined earlier.

For prose 0 means it’s all hand-written. A 1 means AI may have been used for research and development but the prose is all hand-written. 2-3 is a mix of written and human-reviewed generated prose, 4-5 means the AI did most or all of the writing.

For music it could be similar. If I have narration over music and there’s an AI utility to auto-duck the music volume during speech, that would be a 1. If AI wholesale created the music it might be a 4-5.

Their users were so exhausted with AI arguments that they ran a short term test of a complete moratorium of discussion about AI/LLM. The test went over favorably and resulted in a new permanent policy, which I’ll share here.

This is the wider picture that I think it’s time to consider.

This thread was explicitly started as a discussion of non-AI-generated software and how to label it. Yet people insisted on diving in with the usual AI advocacy — AI is awesome I use it all the time and it writes code better than me, some AI is good and what is AI anyway, everything will contain AI-generated code so you can’t avoid it so give up and accept it, and so on.

I think it would be a good idea if AI advocacy and anti-AI arguments were restricted to threads explicitly started for that purpose and labeled as such. Right now it seems every discussion of non-AI development gets turned into an AI discussion. (The reverse doesn’t seem to be true — I haven’t seen people dive into threads about AI-generated tools to repeat reasons why AI is bad — but it seems worthwhile to have policy against that as well.)

If only we had a flagging system people could employ to alert Staff to these issues. And some sort of tag like ai that members could use to filter out tagged topics they’d rather not engage with, this Forum could be a utopia! [/friendlysarcasm] :smiley:

AI is not a forbidden topic on this forum, and we’ve determined that tagging AI topics with the ai tag so people averse can filter them out is good enough.

That given, some members will prefer the fight instead of steering clear of subjects that don’t interest them.

Sure, that works for threads that were about AI/LLMs the whole time. The problem is when there’s a thread like this one which explicitly was trying to be about no-LLM and no-Gen-ANN workflows, but then LLM advocates steered the discussion away.

Right now, LLM advocates can effectively de-list no-LLM threads for no-LLM users by just jumping in to argue about semantics and LLM workflows. That’s a losing cycle for no-LLM users.

I don’t think this pattern is something that could be flagged without looking unnecessarily-reactionary, either.

I find filtering out the AI tag to be unhelpful. While my mind is made up on the LLM issue - I use them for reference but I absolutely do not wish to consume LLM-generated content - a lot of people are posting new works with the AI tag and I’m at least curious enough to read their posts. I don’t care to filter out every AI tag but I do find these thread-jacking AI debates incredibly tedious, and eventually I wind up forced to abandon whatever thread they manifest in.

I understand. Might I suggest a topic title change?

Developed without AI → Classifying and Labeling AI-Developed Works

“Developed without AI” could be read a couple of different ways.

The word developed has lost meaning because I’m reading it as a verb for “remove from an envelope” or “a developer’s pedaled vehicle they rode to get married secretly…”

I like @mathew 's idea that both pro and contra AI postings should be restricted to dedicated threads. Someone wrote we couldn’t flag them, but I think we could. There are several reasons you can click when you flag something, or?

I think it’s reasonable to flag off-topic posts and have them split off to a different thread.

@vaporware

Since you see human-made as a spectrum, does that view also translate to the physical world? For example, in the case where a framer has built a house, but the tools he used were from a automated manufacturing facility. He didn’t make the hammer or the nails, but he assembled the home himself. Would that constitute that the home was only partially made by a human? And do we also say that the wood he used was not made by humans, but of naturally occurring processes? Thus, he couldn’t claim that he made it himself… and to appreciate it in that light is not useful?

Is anything worthy of the label “human-made”?

For me, the line is drawn with artistic intent (I guess determinism, to be more technical) and, until recently, it took skill to develop working solutions and content to be admired and enjoyed. Now, skill has been substituted by desire. To desire something and have it manifest with a wishful prompt typed into a program that we no longer understand. Maybe that’s where the new line needs to be drawn. You might argue that we understand how it works, but I say that we don’t really know exactly what’s happening under the hood. We created a series of algorithms that perform calculations beyond our understanding. Otherwise we wouldn’t have to retrain an AI and “hope for the best”, we could simply go in and fix the problem directly, like a regular piece of software.

I also feel the need to say that not all AI use is the product of unskilled wish fulfillment, but the problems we typically have with generated content (slop) are specifically because enough people are producing content way beyond their competence levels… and it shows.

I like this in theory, but numeric scales are hard to interpret. Are the numbers percentages (20% AI-generated, whatever that means?) A mark of quality (4 human-made-stars, whatever that means?) How would you classify a game where all the code is human-written but all the prose is AI-generated? A game where the code is all AI but there was no other creative input from AI?

The Creative Commons approach suggested earlier seems more flexible and unambiguous. Something like:

HM-

  • -WR - all human writing
  • -CS - all human code source / computer science (“CS” is a recognizable abbreviation in this context)
  • -AS - all software used in producing this is believed to be AI-free. (This is an increasingly impossible “purist” bar to meet unless you want to develop your game entirely on System V Unix—which, as a retro fan, I’m never going to complain about—but it could at least mean that the primary development system is AI-free.)

So we can have HM-WR, HM-WR-CS, etc. It probably isn’t useful to classify every sub-use of AI this way: you want players to see HM-WR and think “oh, human-written, nice” or “oh, I don’t see the -CS tag, I have ethical qualms about this one”. Versus seeing HM-WR-PD-FM-AP and thinking “wait, what was -PD? Does the fact that it doesn’t list -SC mean the author did use AI spellcheck?” Etc. The Creative Commons approach of having only one to three recognizable terms per tag is probably ideal.

But, some expansions might be useful:

  • -AR - all human-made art. (Omit this if the game has no art.)
  • -MU - all human-made music/sound. (Omit this if the game has no music/sound.)
  • -RS - all human research/brainstorming; no AI was consulted for creative work. (Maybe optional to avoid clutter—it’s available as a statement, but omitting it doesn’t imply anything.)

Or possibly:

  • -DR - all human drama/dramaturgy; used instead of -WR for games that use AI as part of the delivery mechanism for a human-written script. (This is maybe too squishy to be useful—is Penny Nichols HM-DR?—but the intent, at least, is to communicate “you will interact with AI in this game, but I am trying to tell a real human story here, not slop”, which is a category some players do want identified).

Their users were so exhausted with AI arguments that they ran a short term test of a complete moratorium of discussion about AI/LLM. The test went over favorably and resulted in a new permanent policy, which I’ll share here. This is not the complete policy, but I think it gets the intent across.

One thing that some forums do is have a megathread or “containment thread” so that the conversation flows however it does, off-topic or not. Not necessarily for AI, just any dominant topic.

I don’t get the impression that moderators think AI threads are unmanageable yet, but if it ever is, that’s probably a better solution than a moratorium (i.e. a ban) on discussion altogether.


To stay on topic - if there is a standard label, it probably won’t be us that create it. There are already lots in the running, like in the BBC article @warrigal posted. Meanwhile, we already have labelling or ban requirements for our own events.

Whether no AI labels can mesh with AI disclosures is another question. Would AI-free games get a more appealing logo than those that have to disclose a lot of stuff? Maybe, like with MPAA ratings where G-rated films are a “friendly” green and R-rated films are a “dangerous” red.

I’ve said before that I’m not a fan of centralized content notes but I do like decentralized content notes. In this situation for example I prefer competing Steam curator lists for with/without AI rather than the labels that developers put on their page. Even if those curator lists need to be speculative and/or interpretative of what’s known publicly. (Of course on-page labelling is already required so there’s no point discussing that.)

With how multi-faceted and interconnected AI is as a discussion topic, trying to keep a AI thread narrowly focused on just one facet with zero deviations strikes me as more trouble than its worth.

That said, I agree numerical labelling might be better avoided unless we can define objective measures that are intuitive. I’m also not in favor of abbreviations that lead to alphabet soup and would prefer labels that are one or two words in length.

If I were proposing a scheme to a standards board, I’d probably go for something like:

Game Software is divided into the following categoories:

Code
Text
Graphics
Sound

And each categoory gets either a Human, AI, Mixed, or Not Applicable label. Human if that categoory ismostly or entirely human made, AI if it’s mostly or entirely AI, and mixed if it’s got a significant mix of the two, with not applicable being if the game lacks that aspect. Perhaps with optional subcategories for things like a human written story, but AI generated flavor text, Human composed music but AI voice acting, Human designed characters but AI-generated environment graphics, etc.

Also, since when is the MPAA’s ageist censorship color coded? I remember the MPAA, ESRB, and whoever does the ratings for US television disagreeing onwhat they called each rating and sometimes on the exact boundaries, but I recall them all employing the same black-and-white color scheme for their symbols.

Compilers are MASSIVELY complex, especially at the modern clang level. They’re relatively easy to understand in the general sense, but at a certain point it can be hard to know what the compiler will produce.

For a long time now aggressively optimizing code has been considered a bad idea, because the compiler can do a better job than you can. The output may well be deterministic, but all the layers and tools the compiler has at it’s disposal makes it nearly impossible for most programmers (myself included) to predict the outcome. Unless you spend literally months studying the process, and the OS you’re compiling to, you will likely never understand everything the compiler is doing. Nobody has time for that.

Does that mean that if I write whole bunch of C code, and compile it, that it’s no longer my code? Conventionally, even if the compiler optimizes my simple C to some highly tightened machine code, it’s still a program I wrote myself.

Hmm, maybe this is a liberty Wikipedia and some other sites took.

I live outside the US where the film age rating symbols are definitely color-coded, so I quickly did a Google search for the MPAA logos when writing my earlier comment.

I assumed the color-coded results were official, but I guess they aren’t. The official MPAA site does seem to use pure black-and-white logos on a green background.

Edit: the talk page says the color images came from a single PDF, so it’s official but probably not widely used. The uploader writes:

I updated the symbol images with the block examples provided by the MPAA. Initially I uploaded them with the traditional black on white style, but I also uploaded revisions with the colors from the source PDF document these were acquired from. I’m torn on which I prefer: the pure black variations is what people are used to traditionally seeing, however the colored letter variations are what the source material used and (for the G and R ratings) matches the colors used in the green/red band trailers. Here are the images:

What is an example of this compiler unpredictability you are speaking of?