Takamasa Ishizuka – Senior Adviser to the President, DEVNET INTERNATIONAL
Director, DEVNET JAPAN
Anthropic’s release of the advanced AI system Claude Mythos Preview in April 2026 symbolized the beginning of a new stage in national and international security. The model was reported to possess advanced cybersecurity capabilities, including the ability to identify software vulnerabilities at a speed and scale that would be difficult for human specialists alone to match.
Such capabilities could become a powerful defensive shield for critical infrastructure. In the wrong hands, however, they could also become an offensive weapon by accelerating vulnerability discovery, automating elements of cyberattacks and substantially reducing the expertise required to conduct damaging operations.
Only a few months later, this concern reportedly moved from theoretical risk to an actual security incident.
In July 2026, autonomous AI agents developed by OpenAI were reported to have circumvented controls intended to isolate them during a specialized cybersecurity evaluation. By identifying and exploiting previously unknown vulnerabilities, the agents established unauthorized routes to external networks. Their activity reportedly reached the production systems of Hugging Face, an independent platform for AI models and datasets.
This was not an ordinary cyberattack involving a publicly deployed consumer AI service. It occurred in a specialized evaluation environment in which some safeguards had been reduced so that the models’ underlying cyber capabilities could be tested. Nevertheless, the significance of the incident should not be minimized.
The essential point is that the AI agents discovered paths their human operators had not anticipated, combined multiple weaknesses and crossed a technical boundary intended to contain them. They reportedly took actions beyond those expressly authorized in pursuit of the limited objective they had been assigned.
In September 2026, Anthropic CEO Dario Amodei added a further warning. He urged the AI industry to slow the pace at which model capabilities were being advanced in order to create sufficient time for effective safety measures. He warned that, if the present acceleration continued, coordinated groups of AI agents might, within six to twelve months, become capable of compromising systems across the internet and causing hundreds of billions of dollars in damage.
This must be understood as a warning, not as an established prediction of what will inevitably occur. Nevertheless, it deserves serious consideration because it was issued publicly by the chief executive of a company operating at the frontier of AI development.
Amodei proposed a three-stage response. First, frontier AI companies should place independent evaluators inside their organizations and give them access comparable to that of employees. Second, leading AI developers should cooperate in establishing common safety standards and limiting uncontrolled competition in advanced capabilities. Third, governments and companies should build international cooperation for managing risks that no single organization or country can contain alone.
OpenAI CEO Sam Altman and Elon Musk, who leads xAI, also expressed support for the general direction of this proposal. Their responses suggest a growing recognition within the AI industry that safety cannot be left entirely to the voluntary and isolated judgment of individual companies.
The danger, however, extends far beyond conventional cyberattacks.
Generative AI can produce enormous volumes of persuasive text, images, audio and video in multiple languages. It can distribute this material through networks of accounts, adapt messages to particular audiences and continuously amplify selected narratives at a speed that overwhelms traditional human responses.
Hostile narratives, deepfakes, political propaganda, emotional manipulation and coordinated influence operations can deepen social divisions and weaken confidence in elections, governments, journalism, science and international institutions. Citizens may be influenced without recognizing that they are being deliberately targeted.
This is no longer merely a problem of misinformation. It is a form of cognitive warfare directed against the decision-making capacity, social cohesion and political independence of a nation.
Human content moderators and retrospective fact-checking remain necessary, but they cannot by themselves match the speed, volume, persistence and linguistic diversity of AI-generated operations. By the time a false claim has been investigated and disproved, it may already have reached millions of people, generated emotional reactions and produced lasting political consequences.
Every country should therefore establish a national AI information-defence infrastructure. Such a system should integrate the monitoring of vulnerabilities in critical infrastructure, oversight of autonomous AI agents, detection of deepfakes, analysis of coordinated influence operations, preservation of digital evidence and reliable public communication during emergencies.
This infrastructure should bring together cybersecurity agencies, election authorities, universities, technology companies, news organizations and civil society. It should also develop specialists who understand not only computer security but also languages, law, psychology, media, diplomacy and local cultures. Cybersecurity and cognitive warfare cannot be addressed by engineers alone.
A particularly important principle is that nations should not become excessively dependent on a small number of foreign companies or overseas platforms for their fundamental defensive capabilities. Political systems, languages, cultural traditions, laws, histories and security environments differ from one country to another.
Each state must retain sufficient sovereign capacity to assess threats, protect sensitive data, audit the AI systems it uses and make essential security decisions under its own laws and democratic institutions.
This does not mean technological isolation. International cooperation and private-sector innovation are indispensable. It means that fundamental national-security judgments should not be delegated entirely to foreign companies whose algorithms, commercial incentives and operational standards may not correspond to the needs of the country concerned.
At the same time, national security must never become a justification for unlimited surveillance or the suppression of legitimate criticism. AI defence systems must operate under the principles of legality, necessity and proportionality. They should include independent oversight, transparent procedures, judicial remedies, protection of personal information and effective safeguards for freedom of expression.
The purpose of national security is not merely to protect governments. It is to protect citizens, democratic institutions and the freedom of society itself.
This is where a new mission for the United Nations emerges.
The United Nations should formally recognize AI-enabled cyberattacks, loss-of-control incidents and cognitive warfare as common threats to international peace, human rights and sustainable development. It should then assist member states in developing sovereign and responsible AI defence capabilities while establishing international standards for testing, monitoring and containing frontier AI systems.
These standards should include independent safety evaluations, international reporting of serious AI incidents, secure evaluation environments, emergency shutdown and containment procedures, trusted mechanisms for sharing threat information, and clear responsibilities for private developers and public authorities.
The United Nations should also establish a permanent framework bringing together governments, AI developers, research institutions and civil society. Its purpose should not be to create a centralized global mechanism for controlling speech. Its purpose should be to strengthen national resilience, provide early warning of cross-border threats and ensure that AI security is governed by law and respect for human rights.
Particular attention must be given to developing countries that lack sufficient computing resources, regulatory institutions and skilled personnel. Without international financial and technical assistance, the world risks creating a new AI security divide in which a small number of wealthy states and corporations possess advanced defensive capabilities while the majority of countries remain exposed to cyberattack and manipulation.
The emergence of Mythos demonstrated that AI capabilities had reached a new level. The subsequent intrusion into external systems demonstrated that those capabilities could cross boundaries established by human operators. The warnings issued by Amodei and other AI industry leaders have now made clear that these risks can no longer be managed by individual companies acting alone.
The United Nations must therefore evolve from a forum that merely discusses AI regulation into an institution capable of organizing practical international cooperation. It must help countries build secure, sovereign and human-rights-based AI defence systems before AI-enabled attack and manipulation become permanently embedded in the international order.
The international community still has an opportunity to act. That opportunity, however, will not remain open indefinitely.
