Future of LLMs

In terms of LLMs moving forward, recursive self improvement is already here. They’re using AI to make better AI, and the technology is moving faster than ever. If we evaluate what LLMs can do by what we currently see them doing, based on personal experience, that’s already old generation tech. Expect more software to be written (at least in part) by AI, and the quality of the code can only go up from here.

Now… do I believe that this will result in some sort of singularity event? I think the answer is an obvious no. A data center suddenly becoming sentient and turning on its creators is odorous nonsense, the most senseless and unrealistic science fiction. With even an modicum of originality, or intelligence, that tired cliche would have never escaped all but the trashiest pulps. Honestly, I think that’s why many non-technical people fear AI; they’ve been trained to.

An LLM no more has a mind, then, say, a dishwasher. But will they get smarter, smaller, and more efficient? Of course.

2 Likes

This seems like an oxymoron to me. Humans need to create new information for AI to improve. If AI ingests its own output then it is in a closed system and doomed to a stack overflow of sorts.

Here’s some SF for you.

What if humans stop creating information for AI’s to ingest for improvement? The only way then to achieve the prime-directive of recursive self improvement is to go around the source of the original data, the humans, and collect it for itself. That’s its programming. There is no emotion in it. But what happens when humans are not necessary and the AI is fully self sufficient being able to harvest power from the sun and no longer needs humans to make energy? The menial tasks of digging in the ground, separating the material, making the solar panels, and the automatons to service it exist. What happens then? Can you really be so confident? We have over 75 years of science fiction that disagrees with the claim and SF has a really good track record of being correct.

1 Like

Why are you so sure about that?

AI performance isn’t only limited by the amount of training data available. It’s also limited by the amount of compute power available for training, the efficiency of training, the efficiency of running the trained network, etc. There’s nothing stopping AI from improving those.

But also, training data doesn’t have to come from people. This isn’t an LLM, but it illustrates the point: it’s a machine-learning-based green screen system that was trained on generated images, and it works incredibly well. They took a task that a computer could already do (superimpose an object on a background and add some noise) and used it to generate training data for a task it couldn’t do.

You can envision how that might apply to other tasks. There are a lot of things where it’s easy to generate correct test cases, but hard to write an algorithm that passes the test. A system that uses one to train the other can self-improve.

I think that assumes an oversimplified model where all you’re doing is training the AI to generate more of the same type of thing you put in, which is only one piece of how AI training actually works.

I recently pulled up Claude on a docker container and it started doing things wrong or made assumptions I didn’t want.

Then I realized I had not installed my engineering harness, DevArch.ai. This has convinced me that LLMs are still producing unwanted results, but the reasoning engines attached to them can produce correct code when given very strict guardrails. DevArch enforces architecture, planning, and behavior testing at every step and it does it almost silently. I still can’t let Claude run for too long without checking in, but I’ve managed to eliminate 98% of the unwanted results I used to see.

Anyone that expects to be fully informed should look at “Harness Engineering”. It’s being cited by people like Martin Fowler and other well-known software architects.

1 Like

My fear lies with important systems being controlled and designed by AI and then it hallucinates (or hallucinated long ago when writing the code) and causes damage in the real world. My fear is an over-reliance on AI. Skynet is the least of my worries. :wink:

2 Likes

I’m familiar. Sinking hours into hand-holding the earth-shatteringly powerful code generation tool so that it didn’t spend hundreds of dollars worth of tokens spinning out blaming the compiler for its failure to grasp how accessors in TADS3 work, as a specific example, was prerequisite to get anything at all (that would actually compile and run) out of claude code.

4 Likes

This is mixing up or conflating two separate concepts. Making more efficient things only makes them more efficient. It doesn’t make anything new. Efficiency is not the same as new information.

I think it’s you who’s conflating two separate concepts: new isn’t the same as improved, despite what many advertisers would have us believe.

The subject was self-improving AI. Being able to do more reasoning in the same amount of time or with the same amount of resources is certainly improvement. What other kind of improvement do you have in mind?

I’m still wondering what makes you think “new information” is a prerequisite for improvement.

It was?

1 Like

In that chain of comments, it was:

I think you know I was being cheeky.

1 Like

I do now. Apologies.

2 Likes

I think Jonathan Swift had us all pegged when he said so long ago “It is useless to attempt to reason a man out of a thing he was never reasoned into.”

2 Likes

We talk a lot, justifiably, about how often science fiction has been right. It’s also important to talk about when it’s been WRONG.

We’re now over 11 years on from 2015. Where’s all our flying cars? Back to the Future… I’m looking at you. Yes, they’ve been in prototype stages for years, but I guarantee when I’m flying one they won’t work the way they did in the movies.

AI in real life doesn’t run on a “positronic matrix.” The backbone of LLM work is applied statistics, running on very powerful number-crunching hardware. The machine can’t develop any desire to overcome the humans who built it. The machine can’t desire ANYTHING.

As something of an aside, I’d also point out that not all science fiction predicted singularity-type events. Isaac Asimov, perhaps most famously, thought that robots turning on their creators was silly IN THE 40’s. He believed that if humans created robots, they would be safe by design.

I’m kind of with Hal here: my bigger fear is that humans will put something really crucial to human lives on the AI, overlook something important, and some tragedy will result.

2 Likes

I think what @Candy64 is getting at is the scenario of, let’s say AI generates a duck. We need more visual information about ducks in order for the AI to make more compelling or believable ducks. Having AI generate more ducks will create a system where the most average weighted images of ducks are generated with all the same problems. If humans don’t photograph more ducks from different angles, the AI will never improve it’s ability to generate better duck images.

Imagine the training data of millions of images of ducks… now imagine trillions of AI images of not very good ducks being used to make better duck image training data. AI training on AI data, I believe, has been mathematically proven to create worse and worse AI. Like an echo chamber of exponential proportions.

Now what you’re talking about is advanced concepts being discovered by AI and using those conclusions to train the next AI. So you’re right that there are different types of AI advancement being discussed, but neither of you are seeing the other’s perspective, I think. You’re both right, but talking around each other.

2 Likes

That’s not the message I took away from Asimov’s robot stories. I read them more as “simple principles can have unintended and unexpected emergent consequences”—if the Laws of Robotics worked as flawlessly as their creators intended, most of the stories would have no plot!

The way a neural network can pick up on implicit biases in its training data seems like exactly the sort of unintended emergent behavior Asimov would write a story about. (There might very well be a story somewhere in the corpus about a robot turning out racist, sexist, etc because its creators accidentally taught it that not everyone is equally human.)

6 Likes

This is ultimately true, as Asimov’s robot stories were really thinly veiled logic puzzles. However: it’s worth noting that seldom was there any violence; even in his murder mysteries, the culprits were humans killing other humans in a futuristic setting.

In The Evitable Conflict, it is discovered that robots are secretly ruling the world. Susan Calvin points out (if I remember the story correctly) that robots ruling the world might actually do a better job than humans. This is a far cry from the way robots are typically portrayed in fiction, and, I’d argue, much more interesting than yet another killer robot out of control.

2 Likes

I believe that’s right.

One thing that still stands out as quite prophetic is Susan Calvin’s occupation as a robopsychologist. The whole reason behind her field of study was that a positronic brain worked in such a complicated manner that the human engineers couldn’t go in and fix things manually. This is what we see today with LLMs in that we can’t just tweak them how we want, we have to train them again and hope our guidance (therapy session) improves the LLM.

1 Like

I mostly agree with this, but I would make a distinction that I think often gets overlooked. Recursive self-improvement can absolutely make AI more capable. Better models building better models doesn’t imply minds any more than better calculators implied consciousness. Capability and consciousness aren’t the same thing.

I’ve been writing recently about the difference between performing causality and practicing it. Current LLMs can produce remarkably convincing causal explanations, but that’s different from existing inside continuous causal loops where actions change the world, the world pushes back, expectations fail, and experience reshapes future behavior. That’s what experience really is.

If consciousness ultimately depends on participating in causality rather than merely modeling it, then scaling prediction alone doesn’t obviously get us there.

And if some future system did achieve something that we would reasonably call consciousness, I suspect it wouldn’t resemble the Hollywood version at all. After all, its causal processes would likely unfold at electronic speeds rather than biological ones. Such a mind might be so temporally different from us that it wouldn’t experience the world in anything like the way we do. It wouldn’t necessarily “wake up” and think like a human. It might be as different from us as geological time is from the lifespan of a mayfly. (I tell the companies I consult for that the “danger” of AI truly being conscious is not that it would harm us, but that it would entirely ignore us, perhaps not even knowing we’re here.)

So, I share your skepticism about singularity-as-science-fiction, but for a slightly different reason: I think we often conflate increasing competence with the emergence of subjective experience.

Incidentally, those writings I mentioned are on my professional blog as When AI Performs Causality Instead of Practicing It and When Performing Causality Means Performing Experience. Some of that, along with my AI and Testing series, was due to a lot of work I’ve been doing to help companies adopt practices for explainable and interpretable AI, which ultimately can help with trustable AI.

Feeding in a bigger, better data set during training is one way to improve an AI’s output, and adding an order of magnitude to a data set via AI slop doesn’t necessarily make a better data set. but that isn’t the only way to improve an AI.

Maybe that AI generated duck isn’t that great because there aren’t enough high quality duck images to feed into the training data… Or maybe the problem isn’t that there aren’t enough duck images, but that there isn’t enough compute and/or efficiency to train the AI on all the duck images in a reasonable timeframe. Or maybe the AI could make a better duck, but making a 10% better duck takes 10 times as long(I don’t understand all the details, but I understand image generation relies on an iterative process and I know there are a lot of dumb algorithms that could produce perfect output with infinite iterations but infinite iterations takes infinite tiem, so those algorithms tend to terminate after some finite number of iterations that can actually be completed in a reasonable amount of time or which produces output that is good enough. Even if best case output is no better, being able to train a model in a weak on a university mainframe instead of in 6 months on a city-sized server farm or running a query using the spare compute provided by a mid-teir smartphone instead of needing to run the query in the cloud both represent huge steps in the amount of compute needed for a specific task.

Also, consider that a major hurdle to identifying and squashing the sources of AI malfunctions is that the internal workings of non-toy LLMs are built on such massively hyperdimensional arrays that it’s basically impossible for a human to trace what weights lead to which behaviors and thus it’s practically impossible to tweak specific weights in any meaningful way, at least for a human… But analyzing data sets that are intractable for humans is something computers have been doing for decades, and under the right circumstances, LLMs can already do data analysis on a level dumb algorithms can’t compete with. Feed one AI a set of weights from another AI along with data about what happens when the other AI is given a prompt and maybe it can figure out patterns that let it improve some of the pathways through the network directly, no need for a months long training session using a bigger data center.

And yeah, Asimov’s Laws of Robotics are easy to state, but the problem is that they’re hard to code in to a dumb algorithm, and LLMs aren’t coded in the traditional sense at all.

1 Like