-- Dario Amodei...is there a kind of threshold moment when the AIs are smart enough to train other genius-level AIs, so whoever gets there first has an unassailable lead?
Despite the presence of the question mark, this sentiment, from a man who stands to gain more than most from public belief (correct or otherwise) that the answer to his question is “yes”, has been widely interpreted as a fact.
While much ink has been spilled over “whether this sort of superintelligence is possible or whether it will be achieved and on what timeline”, and some ink has been spilled over “but what is the definition of superintelligence anyway”, remarkably little ink has been spilled on whether B follows from A: “would the actual or hypothetical creation of a recursively self-improving system (A) truly imply an unassailable lead (B)?”
I do not believe that it does, and the case that it does rests on several assumptions that we simply move past - and would not survive casual inspection if we dwelled on them.
Exponential growth = insurmountable lead
Of course! I mean, of course?
⸮ As anyone who completed middle school math can attest, all exponents are the same number!
Wait, no, that’s not right…
Consider the following two equations:
(n + h)²
n³
In fact, there is no value of a “head start”, h, you can give to n-squared growth that will give it such a lead that n-cubed growth never catches up.
The section is headed "Exponential growth," but n² and n³ are polynomial, not exponential. The argument still works with exponentials (2ⁿ⁺ʰ vs. 3ⁿ), but a technical reader may nitpick it.
-- Claude
A machine capable of improving itself faster than a human engineer could will render human engineering irrelevant
Of course! I mean, of course?
⸮ It’s a well-known fact that there is in fact only one semiconductor engineer on Earth, and because that engineer is The Best, there is no need for a second or third engineer. As we all know, this is true of every single engineering team; once you’ve hired your Best Engineer on the Team, there is simply no need for additional hires. They can contribute nothing.
While I expect to see, within my lifetime, a machine that is better than humans at semiconductor design and model optimization … it requires a leap of faith to believe that besting us at those skillsets in general means that there is absolutely nothing left for us to contribute to the craft.
Continuously improving silicon and artificial neural networks will close all remaining gaps between machine and human intelligence
Of course! I mean, of course?
⸮ We absolutely know enough about natural and artificial neural networks to be certain that we’ve unlocked all of the mysteries of the human brain. Or, if we haven’t, faster silicon most certainly will.
There are still quite a few tasks at which biological neural networks beat artificial ones. To assert that better silicon or something that better silicon spits out will close those gaps, without putting forward a mechanism, or even an enumeration of what those tasks are and a diagnosis of why artificial neural networks struggle … again requires a leap of faith.
To be clear, I’m not arguing that A implies not-B. I’m stating that A does not imply B. Silicon might accomplish those tasks … but it certainly doesn’t tautologically follow that it will.
We’ve been down this road before as a civilization - the idea that intelligence can be boiled down to a single number, or set of numbers, which encompasses All There Is to Know, and there cannot possibly be anything of value that falls outside of that taxonomy. It’s an impressive amount of hubris to combine with a belief that you’ll soon be obsolete.
Powerful enough artificial minds will remove all physical constraints
This one actually has gotten some traction. If machines become sufficiently intelligent and are given sufficient power, the only competitive advantage will come from raw materials.
This makes for good sci-fi. And, if all of the above is solved (a big if), it’s even plausible. More likely, though, the belief that this may be the case drives humans to act as if it is - a scary outcome, whether that belief is correct or not.
The alignment problem, again
I’m far from the first person to comment that there’s no guarantee that a sufficiently computationally powerful AI model will do what we want.
Existential risk largely serves as a distraction from the real harms and risks from the technology as it now stands. Police cameras don’t need a smarter or more powerful AI to misidentify suspects; deepfakes have (predictably) barely made a splash in an already saturated misinformation environment. Since this is ultimately a business and tech blog, though, I’ll focus there.
While most of the sci-fi narratives revolve around a machine becoming too smart and too powerful to listen to us, I’m more concerned with the AI that’s not advanced enough to listen to us, being treated otherwise. After all, I use coding assistants all day, every day.
For most practitioners living in the real, non-fictional world, the “alignment problem” is just “when the product doesn’t fucking work”. It generates some babble that plausibly sounds like it’s trying to do the work (or maybe lying, or maybe making excuses), but it’s actually just failing to accomplish the task it’s given.
The existence of phrases like “workslop” and “slop grenades” can also be thought of as an acknowledgement that our chatbots are “misaligned”, not in any nefarious way, but in the sense that they’re tools that are just not up to the job if you think of them as much more than tools.
In just the past week I’ve seen an LLM…
Suggest I commit insurance fraud to test a software workflow where a vendor did not provide a test environment.
Suggest I delete weeks’ worth of science to fix a conflicted database migration.
Argue with me over what day it was.
Repeatedly simplify both code and prose by removing important detail and adding a narrative history of how it simplified things, thus leaving me with a “summary” that was actually longer than the original.
The model most certainly did not “decide of its own will” to do those things; it just didn’t understand what it was saying, in any meaningful sense. It had enough training data to suggest that I simply must test software before shipping it, but not enough sense to cross-reference that with the non-software implications of what it was saying. Similarly for the data deletion suggestion. The date issue was particularly egregious; the model was just working off of context in an old session rather than checking again.
In the past few years we’ve gotten better at building self-correcting systems, but we’re not really there yet in a generalized sense. It’s, again, unclear how the current system of nibbling at the edges of the problem (automated code reviews, automated test agents, adversarial systems, etc.) really gets us there. We can play whack-a-mole with one failure mode after another, but it’s still a lot of work - these systems are missing a je ne sais quoi that lets them introspect and know that they’re making a mistake (though humans are more than capable of this failing, as well, we’re not incapable of getting past it).
An arms race, just like any other
That’s what I expect for the AI industry - software and hardware - and much of the software downstream of it. Just as the nuclear arms race didn’t end when someone first got The Bomb, this one will not end if (or when) an AI model that can design better chips and/or write better models is developed.
Abandoning the race is probably not an option, and over-investing in it to the point of counter-productivity is possible (see also: Tsar Bomba). There will continue to be advances, innovations, steps forward and backwards, and accidents. Resources will matter. Strategy and tactics will matter. Science and engineering will matter.
