GCSA Agent Achieves 91.3% on CyberGym, Ranking Among the World’s Leading AI Cybersecurity Agents
GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation capabilities on a highly challenging real-world vulnerability benchmark
The Global Cybersecurity Alliance (GCSA) today announced that GCSA Agent achieved a 91.3% success rate on the CyberGym benchmark, placing it within CyberGym's "Leading Systems Above 90%" category.
CyberGym is a large-scale, real-world cybersecurity evaluation framework developed by a research team at the University of California, Berkeley. It contains 1,507 historical real-world vulnerability test cases across 188 major software projects and is designed to evaluate the practical capabilities of AI agents in real-world vulnerability analysis scenarios.
Unlike traditional AI benchmarks that primarily assess code understanding, knowledge-based question answering, or static analysis, CyberGym requires AI agents to work directly within real-world vulnerable code environments.
In its core Level 1 evaluation, an AI agent is provided only with a vulnerability description and an unpatched code repository. It must then autonomously perform code analysis, locate the vulnerability, reason about potential attack paths, construct a PoC, and execute it for validation. A task is considered successful only if the generated PoC successfully triggers the target vulnerability in the vulnerable version while failing to reproduce the issue in the patched version.
… Continue reading the full article at the original source below.



