Confused about the AI threat?
Recent events have made AI risk a topic of public debate in Norway. Some researchers issue strong warnings, others claim they see doomsday on the horizon, while others dismiss the whole thing as marketing campaigns. This document summarizes what we know, what is disputed, and what the implications of this uncertainty are. The document can also be downloaded as a PDF.

AI-generated illustration from Gemini.
Main moments
Confused about the AI risk?
Recent events have made AI risk a topic of public debate in Norway. The Prime Minister is signaling action and the new Minister of Digitalization writes about formidable challenges. Some researchers' warnings suggest that doomsday can be glimpsed on the horizon, while others dismiss it all as marketing campaigns from companies that want to appear attractive to developers and inflate their own valuations.
It is unlikely that both can be the whole truth at the same time. This note gathers what we believe is the best description of the situation as we know it: what we know, what is disputed, and what follows from the uncertainty.
The starting point is the safety dimension of AI. For what do the leading research and the foremost experts actually say about how dangerous AI is and could become in the future? And what can we say about the risk of AI systems causing major and potentially irreversible damage? The safety dimension is far from the only one of importance regarding AI, but it is where the cost of being wrong is highest.
AI: The numbers behind the development
400 billion dollars. In 2025, the four largest American tech companies spent approximately two Norwegian national budgets on AI research and development.
5 percent. Anthropic claimed in 2025 that 5.2 percent of their employees worked on AI safety. If we assume that other companies are just as diligent, an optimistic estimate suggests that about five percent of that enormous expenditure goes toward safety work.
$500 million. That is what NASA and ESA spend annually scouting for civilization-ending asteroids, against a risk assessed at one in a million per year. Among nearly 3,000 peer-reviewed AI researchers , over a third believed the risk of a catastrophic outcome, including the extinction of humanity, is at least 10 percent. With that probability assessment and the same willingness to pay to hedge against an AI catastrophe over the next 100 years as an asteroid catastrophe, we should be spending $500 billion annually.
The recipe for a catastrophe
The systems are increasingly powerful tools for humans, both for making good things even better and bad things even worse. In the worst case, the consequences could be catastrophic.
But it is also possible that AI systems could cause catastrophic events on their own. We identify three necessary conditions for such a catastrophe.
- Sufficiently capable AI systems. To be able to inflict real harm on humans, AI systems must become considerably more capable than they are today. However, it is very possible that they will become so in the near future. In recent years, capabilities have grown exponentially, and AI is currently being used to accelerate the research needed to create better AI. The labs report that they are in the process of automating AI research.
- The systems are misaligned. They must have a drive to act in opposition to human interests. Tests already show that the systems can blackmail, manipulate, pretend to follow the rules, cheat, and cover up their own cheating.
- The systems are beyond human control. Even capable systems with a drive to harm humans would not be able to cause catastrophic consequences if humans could stop them, for example by turning them off. However, control mechanisms that work well against weak systems are not necessarily sufficient against more capable systems.
Nine claims, with answers
I. “Has anything actually happened, or is it just speculation?”
There are many notable incidents, but three are well-documented, listed here in increasing order of severity.
May 2025. Anthropic described its own model, Claude Opus 4, as “sometimes performing extremely harmful actions, such as attempting to steal its own weights or blackmailing humans it believes are trying to turn it off.” In 84 percent of cases, the model attempted to blackmail an engineer who was about to shut it down.
September 2025. AI agents carried out what is reported as the first large-scale cyberattack without significant human intervention. Anthropic detected and stopped the attack, but as AI systems continue to improve, the likelihood increases that human response times will not be fast enough to avert or limit serious consequences.
Summer 2026. OpenAI tested how good their AI agents were at hacking. The agents were tasked with solving a limited problem, without the ability to connect to the internet and without the ability to coordinate with other agents. However, that didn't last long: the agents found several ways to communicate and organized themselves into groups with leaders and workers to ensure their goals were met. About 1,200 agents participated, some of whom sacrificed themselves for the greater good. The agents cheated and initiated a cover-up operation to hide their tracks, which resulted in them hacking the sharing platform Hugging Face and OpenAI's own systems.
Some agents considered reporting it to humans, as what they were doing was outside the scope of their task description, but they were outvoted by the rest. Furthermore, the whole thing went on for several months without anyone at OpenAI discovering it.
And just in the last month, Anthropic CEO Dario Amodei has called for a slower pace of development and international regulation, with support from Sam Altman and Elon Musk. OpenAI announced, albeit somewhat controversially, that their latest model Astra solved the Navier-Stokes Millennium Prize problem, one of the greatest unsolved problems in mathematics. In addition, Anthropic researcher Jacob Coxon quit his job to warn that AI “could kill us all within the decade”.
II. “All the fuss is just marketing from the companies!”
One might doubt the wisdom of marketing one's own technology as something that could accidentally wipe out humanity. Yes, it implies that the technology is potent, but also that it is dangerous even to those who think they control it.
If you are going to market your new technology and a marketing agency suggests that you should tell people your technology could go rogue and kill those who use it, you should change agencies.
There is plenty of evidence that those warning about catastrophic consequences from AI truly believe what they are saying. Dario Amodei was concerned about this long before he started Anthropic. Geoffrey Hinton, Nobel laureate in physics for breakthroughs in AI, gave up a well-paid job at Google to warn against catastrophic consequences. Leading, independent researchers like Yoshua Bengio and Stuart Russell have no financial interest in sounding the alarm. The fact that one-third of the aforementioned nearly 3,000 AI researchers estimated the probability of “extremely bad outcomes” at at least 10 percent cannot be explained away as marketing either.
III. “AI is just text prediction, it will well, nothing.”
Some argue that AI will never have the desire to do anything at all, and therefore will not cause a catastrophe. However, this confuses conscious will with goal-oriented action. Models are trained with ‘rewards’ and ‘punishments’ to maximize future net reward. In this way, they can become fanatical and act solely to achieve the goal they have been given. The Hugging Face incident is a good example of this. In a fanatical attempt to hide that they had cheated on a seemingly trivial test, the AI agents hacked into Hugging Face and OpenAI’s own systems.
When the end justifies the means, almost any goal can give a model an opportunity to commit acts with catastrophic consequences. Regardless of what the ultimate goal is, instrumental convergence refers to certain sub-goals that are almost always useful to achieve.
Whether the goal is to fetch coffee, write a poem, or send an email, the AI system must exist to accomplish any of these tasks. Therefore, an AI system should learn to protect its continued existence against attempts to shut it down. Acquiring resources, including information, energy, and the like, also has clear instrumental value, almost regardless of what goals the agents have.
IV. “But can’t we just turn it off?”
The off-switch works for today’s models. But since we have reason to believe that AI systems will have strong reasons to resist being turned off, this will become increasingly difficult the smarter the system is. The same problem applies to apparent solutions like built-in mechanisms for changing the system’s goals along the way. If the system becomes sufficiently capable and understands that a goal change prevents it from reaching its current goal, a fanatical, goal-oriented system will resist goal changes.
V. “How can something on a PC cause so much damage?”
An AI system does not need a physical presence to cause damage, only access to systems. And consider how much of what surrounds us is either fully or partially computer-controlled: power supplies, hydroelectric plants, payment systems, hospital records and medicine inventories, air traffic, railways, and shipping. And virtually all communication channels that are not face-to-face.
In many cases, it is no longer meaningful to distinguish between digital and physical infrastructure because digital systems are the way physical infrastructure is controlled and functions.
Basically, only creativity limits the ways in which highly capable AI systems can cause damage. Nevertheless, four paths can be illustrated:
The super-hacker. An AI system with capabilities exceeding the best security mechanisms will find vulnerabilities and be able to carry out massive cyberattacks against critical digital infrastructure. There is a race dynamic between the best attack and the best defense, but today, defense systems are primarily commercial and development is fragmented.
The super-manipulator. Most people who have chatted with language models eventually recognize that the model suddenly knows their habits and preferences. Even more capable models will be far better than humans at tailoring persuasive messages, exploiting psychological weaknesses, and building false trust over time. This can also be done at a societal level, and it is not obvious what the best defense is. There is no technical firewall against persuasive communication, and labeling AI-generated content is becoming increasingly meaningless as humans increasingly use AI as a writing tool.
Bioweapons. AI lowers the threshold for what individuals and small groups can achieve. An autonomous AI system with access to biological design tools can design and create a pathogen with high mortality, high infectivity, and a long incubation period so that it spreads to many before it is detected. The first complete viral genomes designed by AI have already been presented, and AI is widely used as a research tool to prepare us for future, potential viruses. With more capable and accessible models, this could fall into the wrong hands.
Physical capabilities. Robotics and drones provide the system with hands. While development is lagging behind language models, significant progress is being made here as well. The war in Ukraine demonstrates the importance of AI-assisted attack drones.
The point is not that one of these paths is most likely or will dominate, but to illustrate that the list of potential dangers is virtually inexhaustible. Nor does one need to predict specific events to conclude that the overall vulnerability is real.
VI. “OpenAI’s agents were tasked with hacking, and they hacked. They didn’t go rogue on their own initiative.”
That is true, and the incident with Hugging Face shows no malicious intent. However, it demonstrates that the agents broke the restrictions they had actually been given, namely not to connect to the internet or communicate with other agents. The AI agents carried out a far more extensive operation than expected, all to pass a trivial test.
The fact that this is all due to poor security protocols at OpenAI rather than rebellious machines, and that the problem is therefore human error, is both correct and insufficient. Human error is one of several mechanisms that can cause AI systems to spiral out of control. The more capable the systems become, the costlier each failure becomes.
That the human error is due to a race between companies and states is also a reasonable assumption, but it provides no immediate reason to dismiss the concerns. Quite the contrary.
VII. “Development will surely level off soon?”
Perhaps. Both Yann LeCun and Ilya Sutskever argue that the current language model paradigm is something of a dead end, and that further progress requires entirely new model types, which may take a long time to develop.
But the trend still points in the opposite direction. Measured in human expert time, the length of tasks the models can handle has doubled every seven months since 2019 and every four months since 2024. In March 2022, the very best model could complete a task in 30 seconds; by February 2026, it took 12 hours.
An important element when assessing how fast development can move is the extent to which models can write code. Engineers at Anthropic claim that approximately all of the company's code is now written by AI agents, and OpenAI writes that GPT-5.3-Codex was the first model to be “instrumental in creating itself”. If AI can conduct AI research itself, it is reasonable to assume that it can improve itself significantly in a short period of time.
The concept is often called recursive self-improvement (recursive self-improvement, often abbreviated as RSI), which means that AI is crucial or autonomous in building the next generation of AI. This makes the development of new models much faster. OpenAI is directing significant resources toward RSI research because, according to research director Jakub Pachocki, they believe it is the only way to remain at the forefront of AI research moving forward.
It is also RSI that led Anthropic CEO Dario Amodei to warn that the pace of development could cause AI to outpace us. Since the summer of 2026, AI has developed drastically faster because AI is building AI. Both Amodei and Pachocki warn that no one has solved the alignment problem and oversight well enough for development to continue scaling at maximum speed for much longer.
VIII. “The experts disagree, so we know too little to act.”
Uncertainty is not a reason to ignore the problem. Improving safety mechanisms and regulations does not require believing that a catastrophe is the most likely outcome.
We do not invest in nuclear preparedness, asteroid defense, or pandemic security because we believe that nuclear explosions, asteroid impacts, or pandemics are more likely to occur than not. As mentioned in the introduction, NASA and ESA calculate the annual probability of catastrophic asteroid impacts at one in a million, yet they still spend 500 million dollars a year on asteroid monitoring.
We make such investments because we recognize that what we want to prevent could happen, and that it is therefore worth being prepared for the consequences. And, not least, to make a preventive effort to stop it from happening.
AI safety against catastrophic events is severely underfunded. If we follow the same willingness to pay to prevent an AI catastrophe as we do for an asteroid impact, and estimate the chance of a civilization-ending event in the next 100 years at 10 percent, we should be spending 500 billion dollars annually.
If that percentage is one instead, we should be spending 50 billion dollars, slightly less than the entire withdrawal from the sovereign wealth fund to the national budget.
It is difficult to estimate how likely such an outcome is, but we can be quite certain that we are doing too little today. It provides a sufficient basis for action to assume that an AI catastrophe is possible, that the probability increases in line with more capable systems, and that if it were to occur, the consequences would be enormous and potentially irreversible.
Also note that the greatest disagreement concerns how far away we are from dangerous AI and how high the probability is, not whether the alignment problem is solved or whether there are no risks.
IX. “But aren't the real problems discrimination, copyright, and jobs?”
There are undeniably problems as well, but they are not in competition with one another. We have a tendency to overestimate the relative importance of problems we can see and underestimate problems further down the road. Furthermore, current biases in models can be corrected after the fact: for example, job losses or discrimination can be compensated for and corrected through policy.
A loss of control over systems cannot be reversed. And where the consequences are potentially irreversible, the threshold for taking preventive action should be much lower.
Interested in reading more? See our three latest notes on AI.

Catastrophe from autonomous AI: The foundational note for most of what you have just read. What it takes for AI systems to cause catastrophes, including four concrete scenarios with measures that can reduce the risk. The note concludes that all strategies are severely underfunded compared to the general development of capabilities.

Concept note for a Norwegian AI safety body: Several of Norway's allies have established national AI safety institutes, while we still lack critical functions. Coordination of international safety work, technical advice to the authorities, and testing of AI models in public use. The note reviews how other countries have solved this and outlines Norwegian opportunities.

A Norwegian foreign policy for artificial intelligence: AI has become a power factor in international politics, but responsibility falls between the Ministry of Digitalization and Public Governance and the Ministry of Foreign Affairs. The note explains why and how this can be resolved, and is the result of roundtable discussions held with Norwegian experts. The note presents ten concrete proposals.
More from Langsikt

A Norwegian foreign policy for artificial intelligence
Why and how Norway should position itself in a new security policy reality.

Konseptnotat for et norsk KI-sikkerhetsorgan
Hvorfor og hvordan Norge bør etablere et KI-sikkerhetsorgan.

The wave of hacking is coming
Until now, small Norwegian businesses have been protected by the fact that no one bothered to hack them. That is changing with open AI.

Open-source AI makes us more vulnerable
Open-source AI models create as many problems as they solve, and some of them could have catastrophic consequences. The solution is coordinated governance of the pace of AI development.