Beyond the Benchmark: A Sociotechnical Perspective on Governance, Risk, and Compliance for Trustworthy AI in Cybersecurity Operations

Authors

  • Abdul Rafay Department of Computer Science, Air University, Islamabad.
  • Anique Ahmad Department of Electrical & Computer Engineering, Air University, Islamabad.
  • Muhammad Saad Department of Computer Science, Air University, Islamabad.
  • Abdullah Saleem Department of Computer Science, Air University, Islamabad.

Abstract

Benchmark accuracy is the most commonly reported evidence for artificial intelligence (AI) systems used in cybersecurity operations, yet accuracy on a fixed test set does not, by itself, establish that a system is safe, governable, or fit for deployment. This paper offers a conceptual, narrative synthesis of the sociotechnical conditions that benchmark-centric evaluation leaves unaddressed: adversarial robustness under realistic threat models, meaningful human oversight, data protection and fairness, and organizational governance, risk, and compliance (GRC). Drawing on a purposively selected set of primary technical standards, regulatory instruments, and peer-reviewed literature and explicitly distinguishing empirical findings, requirements imposed by normative instruments, and the authors’ own proposals  the paper develops three conceptual artifacts: (1) a four-layer sociotechnical evidence profile (technical, human, data-and-rights, organizational); (2) a lifecycle crosswalk relating that evidence profile to the NIST AI Risk Management Framework, MITRE ATLAS, ISO/IEC 42001, and the EU AI Act; and (3) a five-level maturity model for organizational AI GRC capability, illustrated with a worked application to a hypothetical AI-enabled security-operations-center (SOC) system. That worked application is a hypothetical, formative illustration of how the maturity model would be used, not an empirical validation of it, and its provisional finding restricted, analyst-reviewed decision support rather than expanded autonomy is reported accordingly. Each artifact is presented as an unvalidated design proposal rather than a measurement instrument, and each normative claim in the paper is scoped to stated assumptions about consequence, autonomy, and reversibility rather than asserted as a universal criterion of trustworthiness. The paper closes with an explicit agenda for empirical validation, including inter-rater reliability testing, worked case studies, and comparison against incident and audit outcomes.

Editorial positioning. This manuscript should be evaluated as a conceptual framework or perspective article informed by a purposive narrative synthesis. It does not claim systematic-review completeness, legal advice, or validation of the proposed assessment artifacts.

Keywords: Trustworthy AI; Cybersecurity; Governance, Risk, Compliance; Maturity Model; Human Oversight.

Downloads

Published

2026-03-30

How to Cite

Abdul Rafay, Anique Ahmad, Muhammad Saad, & Abdullah Saleem. (2026). Beyond the Benchmark: A Sociotechnical Perspective on Governance, Risk, and Compliance for Trustworthy AI in Cybersecurity Operations. `, 5(01), 6848–6875. Retrieved from https://assajournal.com/index.php/36/article/view/2171