The overlooked divide in the AI debate
Why the deepest disagreement may not be between AI optimists and AI critics, but between two fundamentally different views of what makes AI dangerous
In January I offered here on this Substack a translation into English of my first essay in the Swedish magazine Opulens. Today I have another essay in the same magazine that I want to show you. The title given by the editors translates literally as AI — its runaway development must be stopped, which is close to a repeat of the title given in January, but the new essay actually tackles a more subtle issue haunting contemporary AI debate. Here comes:
The public debate about AI is a concern for all of us, yet it can easily seem disorganized and difficult to navigate, not only for outsiders, but even for those of us who are immersed in it and take an active part in it. A tempting simplification is to divide the participants into, on the one hand, those who view AI’s societal impact primarily as a positive force and want to accelerate its development, and on the other hand, critics who focus mainly on the technology’s risks and therefore advocate a more cautious approach.
The reality, however, is more complicated. Most obviously, the vast majority of commentators recognize that AI offers both opportunities and dangers. Few consistently argue either for flooring the accelerator or slamming on the brakes. Instead, they emphasize the importance of steering the technology wisely so that society can reap its benefits while minimizing its downsides.
Less obvious — but for that very reason more deceptive — is another, largely unspoken disagreement among those of us who primarily emphasize AI’s risks. I count myself in this category, even though I also believe AI has enormous potential to benefit both individuals and society if developed responsibly and accompanied by appropriate regulation. The divide concerns why AI is dangerous. Is it dangerous primarily because it is not intelligent enough, or because it is rapidly becoming too intelligent? The gulf between these two perspectives is often so deep that it makes sense to think of them as two distinct camps.
The first camp focuses on issues such as algorithmic bias when AI systems are used to evaluate people (such as in decisions about bank loans or job interviews), or on their tendency to hallucinate, or on their role in degrading human language. Large language models like ChatGPT and Claude are often dismissed with labels like “stochastic parrots” or “glorified autocomplete.” What these critics rarely take seriously is the extraordinary pace of AI progress, and the likelihood that future systems will be dramatically more capable than those we have today.
The second camp — which is where I would place myself — takes the increasingly steep trajectory of AI progress far more seriously. This does not mean dismissing the concerns emphasized by the first camp, although some of those problems may prove temporary as AI systems become more capable, such as by becoming better at avoiding hallucinations. At the same time, however, a new set of problems is coming into view. One concerns the labor market, as AI systems increasingly surpass human performance across a growing range of occupations. Beyond that lies an even more fundamental question: what happens once AI becomes capable enough to challenge humanity for control? And what would it take for our species to survive such a transformation?
Questions of this kind go back at least to pioneers such as Alan Turing and Norbert Wiener in the middle of the twentieth century. For decades, however, they were largely ignored. Partly this reflected a reluctance among AI researchers to speculate too boldly about future technological advances, but it was also caused by the impression that scenarios in which AI would match or surpass human general intelligence belonged to a very distant future.
Within what I have called the first camp, these questions are still largely ignored, despite the fact that the situation today is profoundly different — for reasons I will return to shortly. When such scenarios are discussed at all, they are typically dismissed as “speculation” or “science fiction.” Rejecting speculation in this sweeping manner reflects a failure to appreciate that any discussion of the future necessarily involves some degree of speculation. As for the science fiction accusation, let me quote from an email I recently wrote to a Swedish computer science professor who criticized me on precisely those grounds:
Your impression that my arguments seem “inspired by science fiction” is largely a consequence of the fact that academic computer science has, for decades, been far less willing than Hollywood to engage with the important question of where AI development might ultimately be heading. [...] To blame me for this regrettable state of affairs is entirely misplaced, because, in fact, the blame for why things have turned out in this way lies squarely on you and your fellow computer scientists.
Those may be harsh words — but I stand by them.
One of the most influential ideas in discussions of whether AI could eventually attain superintelligence — i.e., intelligence vastly exceeding our own across all relevant cognitive domains — is that of recursive self-improvement. The idea can be traced back to mathematicians such as I.J. Good and Ray Solomonoff in the 1960s and 1980s, respectively. It suggests that once AI development is driven primarily not by human engineers, but by AI systems improving themselves, a powerful positive feedback loop will emerge, potentially accelerating progress so dramatically that terms such as the Singularity or an intelligence explosion become appropriate. More recently, researchers like Daniel Eth and Tom Davidson have refined the mathematical models underlying this idea and grounded them more firmly in empirical evidence.
What an increasing number of insiders and independent experts in and around Silicon Valley now believe is that we are rapidly approaching precisely this tipping point, where the feedback loop begins to gather real momentum. One indication is that AI has become so proficient at programming that a substantial share of the code written at leading AI companies such as Anthropic and OpenAI is now generated by AI itself.
On June 5 this year, Anthropic published a report entitled When AI Builds Itself, describing this development in much greater detail. Its authors argue that the tipping point could arrive within just a few years, and they emphasize the risks of triggering such a self-improvement cycle in a world as unprepared as ours. They also discuss the desirability of establishing institutions capable of coordinating a slowdown in AI development if circumstances require it. Just three days later, OpenAI released a statement expressing broadly similar concerns.
What neither Anthropic nor OpenAI says, however, is that development toward superintelligent AI should be paused as soon as possible — despite headlines around the world, including in Sweden, claiming exactly that. If we focus on Anthropic’s report, its actual position is considerably more restrained. The authors merely argue that “it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology” [italics in original]. They explicitly reject the idea that Anthropic itself should unilaterally pause its work, pointing out that doing so would simply allow competitors with weaker safety standards to catch up. Consequently, Anthropic intends for the time being to keep pushing ahead, and the same is true of OpenAI.
The media’s coverage of Anthropic’s report has failed in more ways than one. The sensationalist headlines are only part of the problem. Consider, for example, the Swedish technology columnist Björn Jeffrey, who dismissed the report as “today’s cry wolf,” arguing that Anthropic had “issued similar warnings several times over the past three years” — as if three years were an unreasonably long lead time when warning about what could become the single most consequential turning point in the history of human civilization. Jeffrey’s point, of course, is that the warnings are little more than a cynical marketing strategy by Anthropic. He is far from alone in taking that view. And while it is certainly possible that Anthropic has commercial or other strategic reasons for communicating its risk assessments as it does, there is in fact no need for us to resolve that question in order to judge whether the warnings deserve to be taken seriously. The reason is that awareness of where AI development appears to be heading is now widespread throughout the Californian AI ecosystem and beyond. It is shared by a large number of independent researchers and experts, including, in some cases, Nobel laureates. We are therefore in no way dependent on Anthropic’s own credibility in concluding that AI development could place humanity in an extraordinarily dangerous situation within just a few years, one that may threaten not only our future prosperity but our very survival.
If we are to survive the emergence of increasingly powerful AI, I believe it is essential to halt the ongoing race toward superintelligence between Anthropic, OpenAI, and their competitors. Since these companies appear unwilling to take that step voluntarily, intervention by governments and ultimately international agreements will be necessary. Achieving that, however, will require political pressure and the mobilization of the latent public concern about AI that already exists. (That is one of the reasons I write essays like this one.)
But for that to happen, the public must first understand the scale of what could go wrong with AI. One obstacle to that is the systematic tendency of what I have called the first camp to underestimate both the current capabilities of AI and the direction in which those capabilities are heading. Their views are often so deeply entrenched that even when an AI system is released whose danger stems unmistakably from its competence, rather than from any lack of competence, they still cannot resist downplaying that very competence.
I am referring here to Anthropic’s Claude Mythos Preview, which the company unveiled in April but chose not to release to the general public. Instead, access was limited to a small number of trusted cybersecurity firms because the model had become so superhumanly capable at cyberoffense that a general release could have caused widespread disruption. By consistently using dismissive rhetoric about what AI can already do, as well as what it is likely to become capable of doing, the first camp undermines the public education and opinion-building that I believe are essential if we are to change course and move away from the reckless trajectory that could all too easily culminate in a full-scale AI catastrophe. In doing so, they inadvertently play into the hands of the cynical tech billionaires driving the race forward. For that, I believe they bear a grave responsibility.
The situation is particularly unfortunate because the two camps I have described actually have a great deal in common. Representatives of both are often sharply critical of the behavior of the leading AI companies and advocate far-reaching regulation. I have previously argued that this shared ground ought to make it possible to form a united front against the reckless Silicon Valley elite and the AI accelerationists.
The more I have reflected on the matter, however, the less confident I have become that such an alliance is realistic. The differences in how the two camps understand the nature and future trajectory of AI run very deep, and meaningful cooperation can be difficult when there is fundamental disagreement over whether AI’s greatest problem is that it is too stupid or that it is becoming too smart.
The situation would, of course, be different if the first camp had compelling arguments for its view that AI’s limitations will remain fundamental, both now and in the future. Unfortunately, that is not the case. More often than not, no arguments are offered at all. When arguments are presented, they typically amount to the claim that AI is, at bottom, nothing more than a collection of simple, soulless mathematical operations. But the idea that a system composed of simple components cannot exhibit intelligence or other interesting emergent properties is a fallacy. The easiest way to see this is to consider the human brain, which itself consists of nothing more than atoms and elementary particles mechanically moving around and colliding with one another.
The human brain therefore constitutes an existence proof that intelligence can emerge from an arrangement of simple components, each of which is, in isolation, entirely devoid of intelligence or purpose. Inspired by that example, we are now racing to build a new kind of intelligent entity. Eventually, these systems may become so vastly superior to us that we lose control altogether. That is why we must stop before we reach the point of no return.
But when should we stop? A tech optimist might answer: at precisely the last possible moment, just before further progress becomes impossible to halt. The difficulty, however, is that the profound uncertainties surrounding this unprecedented technological frontier make it impossible to identify that moment. Realistically, then, we face only two possibilities: we stop too early, or we stop too late. For my part, I would far rather we stop too early than too late. And given how far down this dangerous road we have already travelled, I believe the time to stop is as soon as possible.

