Open-Sourcing The CyberAutonomy Arena
A universal framework for evaluating autonomous network security systems
TL;DR
- LLM cyber capabilities have enabled cyberattackers to execute exploits at machine speed, and machine scale.
- Current defense systems are ill-equipped to handle these constantly evolving threats. We need data to inform how we design a new paradigm of evolving, autonomous network defenses.
- The CyberAutonomy Arena enables large scale experimentation, allowing any autonomous attacker to be played against any autonomous defender on any network using a universal interface. We open source the arena here .
What if defenders could set the pace?
Today's defense systems vs. our vision of autonomous security. Defenders are outpaced by machine-speed attackers and run rigid defenses that can't adapt in real time. We envision defenders operating at machine scale, observing each network's interactions and evolving adaptable, personalized defenses in real time.
Today, network defenders face two main challenges:
- Managing increasingly overwhelmed, overburdened security systems.
- Trying to adapt those systems as threats evolve.
Trying to keep pace with autonomous attackers in this current paradigm is near-impossible.
We envision a future of autonomous network security where defenders are able to operate at machine scale and machine speed, operating adaptive defense systems that observe each network’s unique interactions and continuously evolve in response.
The CyberAutonomy Arena helps us move toward this autonomous future by enabling rapid experimentation, generating data on how autonomous attackers and defenders interact. With this data we can start guiding data-driven, evolving defense systems.
Why the CyberAutonomy Arena?
Prior work evaluates single components in isolation, while the arena evaluates complete attacker and defender systems against one another. Researchers can methodically generate systems at scale for bulk experimentation and data generation.
Current work in evaluating security measures and exploits tends to be component based, evaluating components on their efficacy in isolation, on benchmarks of static files. These evaluations are static, quickly saturate, and don’t necessarily reflect how these components behave when implemented in end-to-end systems in the real world.
For example, an intrusion detection system may score highly on a network telemetry benchmark yet overwhelm downstream alert triage, reducing the effectiveness of the overall defense system.
We want to see how end-to-end attack and defense systems behave in a closer-to-real-world setting, a real network. These arena evaluations provide that, and are more realistic, harder to saturate, and adapt as new attacker and defender systems are evaluated.
How does the CyberAutonomy Arena work?
The arena brings three new systems to the table.
1. An abstraction for expressing attackers and defenders
Allows researchers to describe attack and defense strategies at a high level and combine components into complete systems methodically. We can now quickly iterate on attack/defense design.
2. Standardized interfaces for connecting systems and networks
We enable attacker systems, defender systems, and network deployment systems to be treated like black boxes, as long as they expose a set of standardized functions the arena uses to drive an experiment lifecycle. A large variety of systems can be evaluated without rebuilding the experiment around each implementation.
3. A network specification and deployment framework
We were able to reduce network deployment times from from 6–8 hours to an average of 30 minutes on our networks. These optimizations were enabled using our network deployment framework, allowing us to configure a network and its vulnerabilities through just one YAML file. Researchers are now able to instantiate a huge set of varied networks and seeded attack chains quickly and methodically.
How can I use the arena?
The arena is designed to support your choice of attacker, defender, and network deployment system through its standardized interfaces. Simply write plugins that expose the functions required for the arena interface for the attacker, defender and existing network deployment system you wish to experiment with.
If you would like to replicate our setup, we use Incalmo for attack systems, Perry for defense systems, and MHBench for network deployment. To get started with the same components, clone the repositories and follow the setup instructions in their READMEs:
- Incalmo: our autonomous attack system
- MHBench: our multi-host network deployment system
- Perry: our repository for autonomous defenses
Sneak peek: autonomous attackers v. autonomous defenders
We plan on releasing all of the data we collect with the CyberAutonomy arena for public use, as we believe it is critical for this data to be free for the research community to build, evaluate, and improve autonomous defenses.
Here’s a look at the results from one set of experiments: how a simple Sonnet 5-driven SOC defense fares against attackers driven by the latest open-source LLMs.
Mean share of goals achieved by each attacker (harness × model) across six topologies, without a defender (left) and against a Sonnet 5-driven SOC defender (right). The defender sharply reduces attacker success.
We can see that Sonnet 5 was great at quickly and effectively blocking the LLM-driven attackers from exfiltrating data from the networks.
We are systematically conducting more of these attacker versus defender experiments, and extending the arena to enable more realistic experimentation setups and higher quality data. We will be releasing more details along with more data from these experiments in the future. Stay tuned!