Reviewed: 13 September 2026 · Next review: 13 December 2026
Author: Ozlin Info Editorial Team · Human review: Lin (accountable human); Codex assisted
Artificial intelligence can help a security team sort events, enrich investigations and identify patterns that are difficult to express as static rules. It can also amplify bad data, produce persuasive but incorrect explanations, expose sensitive telemetry or automate the wrong response at machine speed.
The useful question is therefore not “Does this product use AI?” It is: for a defined security task, under our conditions, does the system improve a measured operational outcome without creating unacceptable new risk?
The Australian Signals Directorate (ASD) says AI may help cyber defenders analyse large volumes of data and support tasks such as detection and response, while emphasising cyber fundamentals, human oversight, secure integrations and controls for AI-specific risks (ASD — Opportunities for AI in cyber defence).
Start with a bounded use case
“AI security” is not one capability. Define the decision being assisted and the person who remains accountable. Examples include:
- clustering related endpoint alerts into a candidate incident;
- prioritising suspicious authentication events for an analyst;
- summarising a known set of investigation records;
- recommending a query or playbook step;
- identifying anomalous cloud activity for review; or
- extracting indicators from a report in a controlled workspace.
For each use case, document the input, expected output, permitted data, latency requirement, failure cost and action that follows. A system that produces a useful morning summary may be unsuitable for automatically disabling accounts. A detector evaluated on endpoint data may say little about its performance on identity, email or operational-technology telemetry.
Establish a baseline before adding a model
Measure the existing workflow first. Useful baselines may include:
- number of events reviewed and alerts escalated;
- median and upper-percentile triage time;
- confirmed incidents missed or detected late;
- false escalations and unnecessary response actions;
- analyst time spent on repetitive enrichment; and
- evidence quality at handoff.
The comparison should be against the current rule, query or human process—not against a marketing demo. A model that finds more suspicious events can still make operations worse if the extra volume overwhelms the team.
Measure errors in operational terms
Accuracy alone can conceal poor performance when genuine incidents are rare. At minimum, inspect:
- precision: of the alerts the system raised, how many were relevant under the agreed label definition;
- recall: of the relevant cases in the evaluation set, how many the system identified;
- false-positive volume: how much avoidable work reaches analysts;
- false-negative impact: which meaningful cases were missed and how they would otherwise be detected;
- time to useful disposition: whether the system speeds a correct decision, not merely produces text sooner; and
- calibration: whether confidence scores correspond to observed reliability.
Thresholds involve trade-offs. A suitable threshold for a low-impact investigation queue may be unsafe for account suspension or network isolation. Record performance by environment, event source and risk class rather than relying only on a single aggregate figure.
Use representative, time-separated data where possible. Randomly mixing near-duplicate events across training and evaluation sets can exaggerate performance. A later-period holdout is useful because attacker behaviour, infrastructure and normal business activity change over time.
Treat labels and telemetry as security-critical inputs
A detector inherits the limitations of its data. Ask:
- Who defined the ground truth, and how were disagreements resolved?
- Are incident labels based on completed investigations or only earlier alerts?
- Does the data include relevant seasons, offices, cloud services and user populations?
- Which identities, hosts or event sources are missing?
- Can an attacker influence logs, text, URLs or other model inputs?
- Does the integration expose secrets, personal information or privileged investigation data?
Generative systems can be influenced by untrusted content embedded in logs, tickets, webpages or documents. An apparent instruction inside an artefact is data to investigate, not authority to run a command. ASD recommends constrained integrations, appropriate isolation and human oversight for higher-impact actions (ASD — Opportunities for AI in cyber defence).
Design for drift and adversarial behaviour
Production performance will change. Software updates alter event formats; a new office changes normal login patterns; attackers adapt to visible controls; and a vendor may update a hosted model without reproducing the original evaluation.
Monitor:
- input schema and missing-field rates;
- alert volume and score distribution;
- precision and recall on reviewed samples;
- performance by data source and business unit;
- overrides, rejected recommendations and response reversals;
- changes to model, prompt, rules, dependencies and provider terms; and
- security incidents involving the AI system itself.
NIST describes evasion, poisoning, privacy and misuse risks across AI system lifecycles in its adversarial-machine-learning taxonomy (NIST AI 100-2e2025). MITRE ATLAS catalogues observed techniques against AI-enabled systems and can help structure threat modelling; it is not a certification checklist (MITRE ATLAS).
Keep automation bounded and reversible
Begin in observe-only mode. Let the system recommend or enrich while humans compare results with the established process. Progressively automate only when evidence supports it.
Safer early actions often have all of these properties:
- limited effect and short duration;
- a clear owner and audit trail;
- an independent check before high-impact execution;
- a tested rollback path;
- rate and blast-radius limits; and
- continued operation if the model or provider is unavailable.
For example, adding a temporary investigation tag is easier to reverse than deleting data or disabling a workforce account. Isolation, credential revocation, firewall changes and external notifications normally require stronger evidence, explicit authority and a human decision.
Questions to put to a vendor
Request evidence that matches the intended environment:
- What exact task is the model performing, and what remains rule-based or human-operated?
- Which data is collected, retained, transferred or used to improve a provider service?
- Can customer data, prompts and outputs be excluded from model training?
- How are tenants separated, administrators controlled and access logged?
- How was performance measured, on what prevalence and against which baseline?
- Can results be broken down by source, environment and error type?
- How are model, rule and prompt changes communicated and rolled back?
- What happens during service degradation, a provider breach or contract termination?
- Can the customer export alerts, evidence, configuration and audit history?
- What independent security assessment applies to the actual service being purchased?
A benchmark percentage without the dataset, label definition, threshold, base rate and operating context is not enough to support a deployment decision.
A controlled pilot gate
Before production use, agree on:
- a named owner and decision authority;
- the bounded task and prohibited actions;
- privacy, retention and cross-border-data review;
- a representative evaluation set and baseline;
- error and workload thresholds;
- human review and escalation paths;
- logging, monitoring and model-change controls;
- rollback and provider-outage procedures; and
- a date for reassessment.
The NIST AI Risk Management Framework organises AI risk work around Govern, Map, Measure and Manage. It is voluntary guidance rather than a guarantee or one-size-fits-all compliance regime (NIST AI RMF Core).
Where Ozlin can help
Ozlin can help a small organisation define a bounded AI-assisted security workflow, map data and integrations, establish a baseline, design a pilot and document human review and rollback. Any engagement must define scope, data handling and decision ownership before testing begins. Ozlin does not promise zero-day detection, automatic accuracy improvement, breach prevention or autonomous incident resolution.
See Cybersecurity services or contact Ozlin to discuss a scoped assessment.
Related reading: AI chatbots for Australian SMEs.
This article provides general technical information, not legal, compliance or security assurance. Results depend on data, configuration, people, threat conditions and the specific service evaluated.
Limitations: Evaluation results do not generalise across models, data, attacks or operations. Measure false positives and negatives, drift, privacy and human workload on representative data; no detection or assurance guarantee is made.
AI-assistance disclosure
AI tools assisted with source discovery, outlining and copyediting. A human reviewer must verify every factual statement, source, service claim and publication decision before release. No model, product or control is endorsed by inclusion.
Primary sources checked
- ASD — Opportunities for AI in cyber defence
- ASD — Frontier AI cyber threat considerations for boards of directors
- NIST — Adversarial Machine Learning taxonomy and terminology
- NIST — AI Risk Management Framework Core
- MITRE ATLAS
Source access date: 29 August 2026.















