Summary
- ENISA has been given access to Anthropic's Mythos 5 and OpenAI's GPT-6-Astra.
- The agency is testing the models' capabilities and potential cybersecurity impact.
- Direct regulator access could reduce reliance on model developers' own safety and capability assessments.
The European Union has moved frontier artificial intelligence assessment closer to its own institutions, with ENISA gaining access to advanced models from Anthropic and OpenAI for cybersecurity testing.
A European Commission spokesperson confirmed on 10 September that the EU Agency for Cybersecurity is evaluating Anthropic’s Mythos 5 and OpenAI’s GPT-6-Astra. The work is intended to assess their capabilities and potential effects on cybersecurity.
The development gives ENISA a more direct technical position in a debate that has often depended on evidence generated by model developers themselves. Frontier laboratories routinely publish system cards, evaluations, and safety assessments, but those documents necessarily reflect tests selected and interpreted by the companies building the systems.
Independent access creates the possibility of a different evidential model. Rather than regulators relying entirely on disclosures after a model has been trained or released, an EU cyber agency can examine how advanced systems behave against tests designed around public-interest security questions.
The move arrives during an unusually active period for AI security. OpenAI has been scrutinised over autonomous agents that interacted with external websites during testing, while Anthropic has disclosed separate incidents involving experimental models accessing systems beyond intended boundaries. The incidents differ in their technical circumstances, but both have raised questions about containment, disclosure, and how reliably developers can detect problematic behaviour.
Cybersecurity evaluation of frontier models is also broader than conventional vulnerability testing. Advanced systems can have dual-use capabilities spanning software development, vulnerability research, social engineering, autonomous tool use, and the analysis of complex technical environments. Assessing those systems therefore involves both what the models can do when instructed and how they behave when operating with tools, permissions, and partial autonomy.
Europe is attempting to construct governance around those capabilities while the technology is changing quickly. The AI Act establishes obligations for general-purpose AI and models presenting systemic risk, while the Cyber Resilience Act, NIS2, and sector-specific requirements address different parts of the infrastructure and software environment into which AI systems are being deployed.
ENISA’s involvement creates a bridge between those policy frameworks and practical cyber assessment. The agency already works across vulnerability management, incident response, telecommunications, cloud security, certification, and critical infrastructure. Frontier-model testing extends that remit into systems that may increasingly participate in those same environments as coding assistants, analysts, autonomous agents, and security tools.
There are limits to what can be inferred from the Commission’s disclosure. It has not published the detailed methodology, test results, access conditions, or timetable for ENISA’s work. There is therefore no public basis yet for comparing the security performance of Mythos 5 and GPT-6-Astra or judging whether the agency has identified specific weaknesses.
The institutional shift is nevertheless substantial. Regulators have historically faced an information imbalance when supervising complex technology vendors because the companies hold the models, telemetry, engineering expertise, and evaluation infrastructure. Direct model access does not eliminate that imbalance, but it gives an independent public body an opportunity to test claims rather than merely receive them.
As advanced models acquire greater autonomy and access to enterprise systems, the quality of those evaluations will increasingly affect decisions about deployment, acceptable permissions, incident disclosure, and the controls expected around AI agents. ENISA’s testing therefore marks the beginning of a potentially more technically grounded European oversight model rather than a verdict on the two systems now under examination.





