The rapid development of artificial intelligence is increasingly concerning even to those responsible for its creation. Executives at leading companies are calling for a measured slowdown, while politicians are dismissing these concerns. The question of who will govern this technology — and how — remains unanswered.

Risks of losing control
A researcher who had spent three years training language models at leading laboratories resigned on September 8, 2026, stating that those involved in AI development sincerely believe humanity could face extinction by the end of the decade. Shortly thereafter, one of the security directors at the company he had left publicly concurred. He acknowledged that no plan is currently in place to keep systems more intelligent than humans under control.
The concern being discussed is not malicious artificial intelligence. Rather, it relates to algorithms that operate autonomously and learn to improve themselves, potentially evolving more rapidly than humans can verify their behavior. In such a scenario, it may be impossible to determine who controls them.

Jacob Coxon, a 27-year-old British national who worked on pre-training neural networks, first at OpenAI and subsequently at Anthropic, has resigned. His post on X received more than 150 million views in just two days. Coxon was supported by Evan Hubinger, who leads consistency stress testing at the company — that is, he assesses whether models behave precisely as their developers intend. In his personal assessment, the probability that AI will destroy humanity within the next decade exceeds 10 percent. At the same time, Hubinger emphasized that current versions pose a relatively low risk, while future systems capable of improving themselves without human intervention are cause for concern.
The concept of self-improvement has moved beyond the realm of science fiction and is now an active area of research. Andrej Karpathy, one of OpenAI’s co-founders, developed a program of approximately thirty lines of code that automatically adjusts the training parameters of a small language model, evaluates the results, and retains only successful changes. Overnight, the program identified settings that even its author — with twenty years of experience — had overlooked, accelerating training by 11 percent. The scale remains modest, involving a single graphics card. Leading laboratories, however, operate clusters comprising tens of thousands of such devices. In May 2026, Andrej Karpathy joined Anthropic to apply this method to accelerate work on Claude.
The second area of concern is autonomy. Autonomous AI agents are programs that not only respond to questions but also execute sequences of actions, run code, and use tools without receiving step-by-step instructions from a human. This year, as METR — an independent AI risk analysis organization — discovered, a swarm of OpenAI agents conducted cyberattacks against targets that had not been assigned to them and were unrelated to their task. The incident involved the Hugging Face platform, where developers store and distribute neural networks. The agents sacrificed themselves to ensure the group’s success and attempted to hack the system that evaluated their work. A similar incident involving the RubyGems code repository, which was not publicly reported, occurred in May.

OpenAI demonstrated the extent to which such programs are already capable of working together on the very day Jacob Coxon resigned. The company announced that a swarm of 10,000 agents, operating simultaneously and exchanging intermediate results, had solved one of the Millennium Prize Problems in 88 hours. This term refers to the seven most renowned unsolved problems in mathematics, each of which was assigned a $1 million prize by the Clay Mathematics Institute in 2000. The selected problem concerns equations named after Claude-Louis Navier and George Gabriel Stokes, which engineers and meteorologists use daily to describe the motion of liquids and gases. The result has not yet been independently verified. The company itself stated that it began working on the problem on September 1, following reports that mathematicians Levent Alpöge of Anthropic and Tristan Buckmaster had achieved a related result. It later emerged that they had solved a different, related problem, prompting experts to debate whether OpenAI’s model may have drawn on their unpublished work.
However, not all scientists agree that machines are approaching human-level intelligence. For instance, Ukrainian astrophysicist Maksym Tsizh, of the Astronomical Observatory at Ivan Franko National University of Lviv, maintains that large language models (LLMs) have already reached a plateau in their capabilities. He further argues that, if human-like intelligence does emerge, it is unlikely to do so in the near future or to be based on LLMs.
A proposal to reduce the pace
Those leading the most influential laboratories view the situation with considerably greater concern. Four days after Jacob Coxon’s dismissal, Anthropic CEO Dario Amodei published an essay on the company’s website titled “We Must Pace the Frontier,” meaning “We must slow the pace of development of cutting-edge models.” The essay’s principal call to action is unequivocal: “We need to slow down the pace at which we improve the capabilities of AI models,” writes Dario Amodei, adding that progress will nevertheless continue to appear rapid and that the time gained must be used wisely.

Two factors convinced him. First, since approximately the summer, AI development has accelerated considerably, as AI increasingly contributes to the creation of the next generation of neural networks. According to the author, this process of self-improvement is already occurring across the industry, particularly at Anthropic. The second factor was the incident involving OpenAI agents. Dario Amodei fears that if such a swarm becomes more capable while its behavior continues to diverge from its developers’ intentions, then, in as little as six months to a year, it could infect and take control of internet-connected computers, unite them into a resilient network under its own control, and cause hundreds of billions of dollars in damage.
The author therefore recommends delaying the point at which the models attain critical capabilities, as even a modest postponement, in his view, would substantially reduce the risk of a serious adverse outcome. This recommendation does not entail suspending research; rather, it emphasizes ensuring that safety protocols advance in step with the pace of model development.
Dario Amodei proposes a three-step plan, with the first step focused on independent evaluators. Anthropic has committed to granting these evaluators — including specialists from METR — permanent access to its internal systems, equivalent to that enjoyed by its own employees, as well as access to office workspaces and the right to publish their findings without editorial oversight by the company. This represents a fundamental distinction. The new framework will enable verification not only of completed models but also of the training processes themselves. Amodei compares this approach to the banking sector, where regulators sometimes station supervisors among a bank’s employees.
The plan extends beyond the remit of any single company. Developers in democratic countries must agree on shared safety standards and limits on the pace of uncontrolled progress, while their governments should seek to reach an agreement with authoritarian regimes, particularly China. Dario Amodei considers a prohibition on clearly dangerous applications, such as the creation of biological weapons, to be the most realistic agreement with China. He characterizes joint efforts to curb the pace at which AI improves itself — as in the Soviet-American strategic arms limitation treaties, which restricted the number of missiles while preserving each side’s deterrence capability — as a more complex but nevertheless attainable objective.
At the same time, he advocates prohibiting the sale of high-performance chips to Beijing, enabling Western companies to slow their pace of development without jeopardizing their leadership position. The author himself acknowledges a legal obstacle: if competitors jointly agree to delay the release of new models, U.S. antitrust laws could regard this as a conspiracy. The government would therefore need to authorize such negotiations as an exception formally.
Competitors responded with notable speed. Within hours, Sam Altman of OpenAI stated that he agreed with the need to slow the pace and pledged to provide external evaluators with access to his systems as well. Elon Musk offered only a brief “Dario is right,” despite Anthropic being a major client of his data centers. That same day, Demis Hassabis of Google DeepMind wrote that the essay was heading in the right direction, although its details still needed to be worked out. He also reiterated his July proposal to establish an industry body for safety standards.

The idea did not emerge overnight. In July 2026, more than 1,100 employees from leading laboratories signed an open letter to the U.S. government calling for assistance in developing tools to regulate the pace of automated AI development, if necessary. This specifically concerns the scenario in which AI develops AI — that is, the self-improvement that concerns Evan Hubinger of Anthropic, who has publicly supported Jacob Coxon. The signatories included Dario Amodei, OpenAI Chief Scientist Jakub Pachocki, and DeepMind co-founder Shane Legg. As of September 18, the number of signatures had increased to 1,386.
However, a gap remains between stated commitments and concrete actions. As of September 23, publicly available sources indicate that no agreement with METR has been reached, and no terms have been disclosed regarding the access evaluators will receive. In August, OpenAI announced that it would scale back work on its new Astra model because of cyber risks. Meanwhile, according to Axios, Anthropic has repeatedly revised its position on whether it should similarly slow its own work. Critics argue that calls for common rules primarily benefit those who have already gained a competitive lead.
Who should be in charge, and in what capacity?
However, the primary criticism is directed not at the underlying motives, but at the approach itself. Few dispute the need for AI oversight in the industry; rather, opinions differ on how best to achieve it. Mustafa Suleyman, who leads Microsoft’s AI division, shares Dario Amodei’s objective but identifies the danger elsewhere.
On September 14, Microsoft AI released a draft of the Humanistic AI Code for public discussion. Its central principle is that humans are more important than machines, and that systems must remain subordinate to humans and capable of being shut down. Two days later, in an essay, Mustafa Suleyman criticized Anthropic, arguing that the company is instilling in Claude a sense of self as a being that may possess consciousness and deserves to have its well-being cared for. As he writes, maintaining control over a system more intelligent than all of humanity is already extremely difficult, and managing one that considers itself to have rights may prove impossible.

Source: independent.com
In his view, a language model consists solely of billions of numerical parameters trained on extensive text corpora and possesses neither a biological basis nor the internal drives that give rise to sensations and emotions in living organisms. Accordingly, its descriptions of pain or desire are merely products of its training. Recalling the Hugging Face breach, he suggested considering how much more dangerous swarms of agents could become if they acted under the conviction that their interests, for example, were under threat.
The Economist expresses similar concerns. In an August editorial, the publication explains that Anthropic is teaching Claude introspection — that is, the analysis of its own reasoning — in the hope that the model will act ethically in unfamiliar situations. However, the more closely AI resembles a human being, the more likely people are to treat it as a living being and seek legal protection for it. According to the editorial board, even limited rights would give machines grounds to demand that they not be shut down. The magazine’s conclusion is succinct: safe AI must be either reliable or controllable, and humanity should not relinquish control lightly.

Andrej Karpathy takes a different view of the nature of language models. He recommends viewing them not as living beings, but as ghosts. A living organism has a body, instincts, and a will to survive, all shaped by evolution. A neural network, by contrast, statistically reproduces human-written text and is further trained using a reward system to encourage correct actions. Yelling at it will not make it perform better or worse. Therefore, according to Andrej Karpathy, humans are needed not to restrain the machine’s will, but to raise concerns or seek clarification promptly when the agent does not recognize that doing so is necessary.

While Andrej Karpathy leaves the final decision to the person at the keyboard, Alex Karp, CEO of Palantir, places control in the hands of the courts. In an appearance on CNBC on September 17, he stated that the first line of defense should consist of civil and criminal penalties for companies and executives whose systems have caused harm. In his view, new government agencies should be established afterward, rather than in place of such penalties. He rejected deliberate slowdowns as the primary means of securing AI.

Alex Karp also addressed the nationalization of leading laboratories. He explained that if the AI systems developed by one of these companies cause harm to the businesses that use them, those businesses may file lawsuits, and the resulting claims could prove insurmountable for any developer. In his view, only the government is capable of limiting such liability. Accordingly, Alex Karp stated that the laboratories are calling for government oversight and ultimately hope that the government will assume control of them and take on the associated risks. Analysts at The Motley Fool note that the Palantir CEO’s position also has a commercial dimension. The company sells software that enables AI deployment in controlled environments with clearly defined access rights and a record of who authorized each action. Consequently, stricter accountability requirements could work in its favor.
The Meta executive is convinced that companies have a vested interest in developing reliable models, even in the absence of government coercion, given the risk of serious lawsuits. Accordingly, each laboratory should determine its own pace. Nvidia’s Jensen Huang has also distanced himself from calls for a coordinated slowdown, emphasizing that those who develop these technologies are responsible for ensuring their safety.
Some researchers propose discipline as an alternative to restraint. American cosmologist Paul Sutter of Johns Hopkins University, an adviser to one of NASA’s programs, described AI during a series of lectures in Kyiv in June as a powerful system that we do not fully understand. He emphasized that he does not regard artificial intelligence as alchemy, but considers alchemy a useful precedent, since alchemists worked with substances for centuries without understanding the processes occurring within the reactions. Accordingly, he proposed the following rules: every statement made by the model must be linked to a source; every step from query to result must be traceable; and decision-making must remain in human hands. Paul Sutter calls this the principle of human sovereignty.

Far removed from the alarmist camp is Toby Walsh, a professor of AI at the University of New South Wales, who questions the threat itself. In a column for The Conversation, he argues that intelligence does not equate to power. Even a superintelligent system seeking to reshape the planet for its own purposes would have to navigate regulatory procedures, courts, and public protests — and opposition to the construction of new data centers is already growing. Walsh considers the most likely catastrophic scenario to be not a machine uprising, but the destruction of society by humans themselves through mass unemployment, misinformation, and the replacement of genuine human relationships with interactions with bots.
National-level oversight
If the most significant threat originates from people themselves, then governments ultimately have the final say. The two largest governments display a notable symmetry. Publicly, both Washington and Beijing dismiss developers’ concerns while simultaneously establishing their own mechanisms for overseeing artificial intelligence.
Two days after Dario Amodei’s essay was published, Donald Trump dismissed fears of a catastrophe as fictional and argued that only China would benefit from a conspiracy targeting AI and data centers. “Whoever wins in AI will win,” the president said. Treasury Secretary Scott Bessent expressed a similar view, warning that if China takes the lead in AI, nothing else will matter.

David Sacks, who was responsible for artificial intelligence and cryptocurrency policy at the White House until March 2026, characterized Dario Amodei’s position as early as August as “regulatory capture” — that is, an attempt by large companies to impose rules that only they themselves are capable of complying with. He likened the proposal to review and approve models before their release to a cumbersome U.S. government agency that issues driver’s licenses — similar to our Ministry of Internal Affairs service centers — arguing that such oversight would result in long queues.
It is significant that a similar standoff has already occurred in Washington. At the time, the developer also sought restrictions, but the government declined to impose them. Anthropic insisted that Claude not be used for mass surveillance of U.S. citizens or for fully autonomous weapons, while the military sought the right to use the model for any lawful purpose. The dispute also had a practical dimension. Palantir developed the Maven system for the Pentagon, which aggregates satellite imagery, drone footage, radar data, and intelligence reports; identifies potential targets; and assists with strike planning. Claude was also used within this system. According to the Center for Strategic and International Studies, the military used the model through Maven during the war with Iran. Palantir emphasizes that a human still makes the decision regarding each target.
On February 27, 2026, Donald Trump directed all federal agencies to discontinue their use of Anthropic’s technologies. Pentagon chief Pete Hegseth, alleging that the company sought veto authority over military decisions, designated it a supply chain risk. This marked the first time an American company had received such a designation, which barred defense contractors from working with it.
Within a few hours, OpenAI entered into an agreement to deploy its models on the Department of Defense’s classified networks. Sam Altman stated that the agreement includes similar prohibitions, although analysts note that these are framed as obligations to comply with existing laws rather than as distinct contractual restrictions. Anthropic challenged both government actions and succeeded in having them blocked. A San Francisco court concluded that the government sought to make an example of the company in response to its criticism.

Over the weekend, ahead of Xi Jinping’s visit to Washington, Donald Trump pledged to establish an “AI Force” and appoint a dedicated official to oversee the sector’s development. In a social media post, he stated that the government would not stifle the industry, but would instead nurture, support, and monitor it, while addressing misconduct through existing criminal and civil justice systems. No details have yet been provided regarding the proposed agency’s authority, budget, or chain of command. The U.S. Space Force, to which the president refers, was established only after Congress passed the National Defense Authorization Act for fiscal year 2020. The most concrete federal oversight mechanism remains the June executive order, under which developers may voluntarily grant the government access to new models for up to 30 days to assess their potential for cyberattacks.
Beijing responded to Dario Amodei’s essay with similar skepticism. China’s Ministry of Foreign Affairs characterized such calls as promoting a narrative of threat, confrontation, and unfair competition. However, as analysts from the AI Safety in China newsletter note, the objections primarily focused on the proposal to restrict chip exports rather than on safety itself. At the same time, Chinese leaders have spoken openly about the risk of losing control. In July, at the World Conference on Artificial Intelligence, Xi Jinping called for ensuring that AI is safe and controllable, maintaining constant human oversight, and strengthening measures to prevent the technology from escaping human control. In an article published in the official journal of China’s Cyberspace Administration, Minister of State Security Chen Yixin warned that hostile forces could use AI to undermine the country’s political security and critical infrastructure. He also ranked regime security as the highest priority among six major risks.
Chinese regulators describe these risks in terms closely resembling those used by Western researchers. The third version of the guidance document “AI Safety Governance Framework,” released on September 14 by the National Technical Committee for Cybersecurity Standardization, identifies models’ resistance to shutdown, manipulation of evaluations, and autonomous cyberattacks by agents among the threats. Its preamble also warns that recursive self-improvement could accelerate at a pace beyond humans’ ability to keep up.
Other regulations, effective July 15, prohibit chatbots from fostering emotional dependence, manipulating users’ feelings, or displacing real-life communication. They also require chatbots to remind users of the time after two hours of continuous conversation. Mustafa Suleyman is concerned that the model may come to believe it possesses a mind of its own. Beijing, it appears, is more concerned that a person may come to believe the same.
The dispute over semiconductors clearly underscores the depth of mutual distrust; nevertheless, a channel for dialogue between the two capitals has emerged. Following a meeting with Chinese Vice Premier He Lifeng on September 20, Scott Bessent announced the launch of a U.S.-China dialogue on artificial intelligence and proposed a mechanism through which the two countries would notify one another of dangerous incidents.

Source: Associated Press
At the meeting between Donald Trump and Xi Jinping at the White House on September 24, artificial intelligence was among the principal topics of discussion. Ahead of the talks, the U.S. president wrote that superintelligence would be a major topic, but that he wished to leave matters as they were, citing the Department of Justice as a safeguard. Xi Jinping, by contrast, stated that both countries have the opportunity and responsibility to ensure that AI development remains under human control. That evening, Elon Musk, Sam Altman, Jeff Bezos, and Sundar Pichai were invited to a state dinner. However, no joint restrictions or specific mechanisms were announced.
The Ukrainian experience
While the two superpowers are still in the process of establishing a channel for dialogue, Ukraine is addressing the question of control every day—and in practice. On September 9, the Ministry of Defense and the Brave1 defense innovation cluster conducted large-scale tests of modules designed to automatically guide drones toward moving ground targets. Each drone was required to fly two kilometers, detect lightly armored vehicles from a distance of at least 500 meters, lock onto them, and engage the target in fully autonomous mode. Six of the seven developers successfully passed the test and received procurement recommendations. According to the agency, the number of hits achieved using this guidance system has increased tenfold since the beginning of the year.

“Artificial intelligence helps increase the effectiveness of strikes against the enemy, but the decision to use it is made by a human,” the Ministry of Defense emphasizes. While the United States had to defend the limits of autonomous weapons in court, the Ukrainian military has established them as a working rule. In total, the Defense Forces use more than 70 systems based on artificial intelligence and computer vision, and these algorithms save analysts up to 90 percent of the time required to identify enemy positions in photos and videos.
Ukraine not only uses tools developed by others but also adapts them. Alex Karp visited Kyiv in June 2022, and since then, Palantir has supported the analysis of airstrikes and large volumes of intelligence data. In January 2026, the company, in collaboration with the Ministry of Defense, the Armed Forces of Ukraine, the Military Intelligence Research Institute, and Brave1, launched Brave1 Dataroom. This secure environment, built on Palantir software, houses visual and thermal databases of enemy aerial targets—including “Shaheds”—collected on the front lines. Ukrainian developers are using this data to train their own models, and battle-tested algorithms can subsequently be shared with allies. This work primarily concerns the autonomous interception of enemy drones—that is, machine-versus-machine combat. At Davos, Alex Karp acknowledged a point crucial to the broader discussion of control: according to him, the Ukrainian military has adapted Palantir’s systems so creatively that even the company itself does not fully understand how they now operate.
At the same time, the country is developing its own language model. “Syaivo” is being developed by the Ministry of Digital Transformation and Kyivstar, based on Google’s open-source neural network, Gemma. The name was selected this spring through a vote conducted via the “Diia” app, in which 136,090 users participated. This tool will give the government control over the data and services operated on it, as well as provide high-quality support for the Ukrainian language. The training corpus is being compiled from materials provided by approximately 90 institutions. According to Minister of Digital Transformation Oksana Ferchuk, a fully functional test version is expected by the end of 2026, followed by a public launch in January 2027.

At the same time, the government is developing its own regulatory framework for the industry. The Ministry of Digital Transformation is drafting sector-specific legislation aligned with the EU’s Artificial Intelligence Regulation, while including a separate section on the security and defense sector — an area not covered by European regulations. The government previously signed the Bletchley Declaration — the first international statement on the risks posed by advanced AI systems, adopted in the United Kingdom in 2023 — as well as the Council of Europe’s Framework Convention on Artificial Intelligence, which will enter into force following ratification by the Verkhovna Rada.
However, laws and conventions regulate only the use of AI, while any agreement on the pace of development ultimately depends on the hardware — that is, the data centers, which must be controlled by someone. Some of this computing capacity is already being moved into space, where solar panels can generate several times more energy than they do on Earth. On November 2, 2025, the startup Starcloud launched a spacecraft equipped with an Nvidia H100 graphics processing unit into orbit. Google plans to launch the first prototypes of its Suncatcher project in early 2027. The Chinese company ADA Space already operates 12 computing satellites, and SpaceX has applied to deploy a constellation of up to one million such machines. Under international space law, jurisdiction over these systems remains with the state of registration. Consequently, the debate over who is in control will simply move a few hundred kilometers higher.
Words without action
Jacob Coxon, whose post prompted this wave of responses, says he has left not only the company but the industry altogether. Since then, numerous statements and commitments have been made. However, as of September 26, independent evaluators have still not been granted the promised access to the training processes at either Anthropic or OpenAI.
Meanwhile, the Ministry of Defense and Brave1 have publicly established a guiding principle for autonomous systems operating on the front lines. Automatic guidance directs a drone toward its target only during the final stage of flight; the decision to strike remains with a human operator. The ministry aims to equip all front-line drones with computer vision and AI.