Tag: model evaluation

  • AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype

    AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype

    Reviewed: 13 September 2026 · Next review: 13 December 2026
    Author: Ozlin Info Editorial Team · Human review: Lin (accountable human); Codex assisted

    Artificial intelligence can help a security team sort events, enrich investigations and identify patterns that are difficult to express as static rules. It can also amplify bad data, produce persuasive but incorrect explanations, expose sensitive telemetry or automate the wrong response at machine speed.

    The useful question is therefore not “Does this product use AI?” It is: for a defined security task, under our conditions, does the system improve a measured operational outcome without creating unacceptable new risk?

    The Australian Signals Directorate (ASD) says AI may help cyber defenders analyse large volumes of data and support tasks such as detection and response, while emphasising cyber fundamentals, human oversight, secure integrations and controls for AI-specific risks (ASD — Opportunities for AI in cyber defence).

    Start with a bounded use case

    “AI security” is not one capability. Define the decision being assisted and the person who remains accountable. Examples include:

    • clustering related endpoint alerts into a candidate incident;
    • prioritising suspicious authentication events for an analyst;
    • summarising a known set of investigation records;
    • recommending a query or playbook step;
    • identifying anomalous cloud activity for review; or
    • extracting indicators from a report in a controlled workspace.

    For each use case, document the input, expected output, permitted data, latency requirement, failure cost and action that follows. A system that produces a useful morning summary may be unsuitable for automatically disabling accounts. A detector evaluated on endpoint data may say little about its performance on identity, email or operational-technology telemetry.

    Establish a baseline before adding a model

    Measure the existing workflow first. Useful baselines may include:

    • number of events reviewed and alerts escalated;
    • median and upper-percentile triage time;
    • confirmed incidents missed or detected late;
    • false escalations and unnecessary response actions;
    • analyst time spent on repetitive enrichment; and
    • evidence quality at handoff.

    The comparison should be against the current rule, query or human process—not against a marketing demo. A model that finds more suspicious events can still make operations worse if the extra volume overwhelms the team.

    Measure errors in operational terms

    Accuracy alone can conceal poor performance when genuine incidents are rare. At minimum, inspect:

    • precision: of the alerts the system raised, how many were relevant under the agreed label definition;
    • recall: of the relevant cases in the evaluation set, how many the system identified;
    • false-positive volume: how much avoidable work reaches analysts;
    • false-negative impact: which meaningful cases were missed and how they would otherwise be detected;
    • time to useful disposition: whether the system speeds a correct decision, not merely produces text sooner; and
    • calibration: whether confidence scores correspond to observed reliability.

    Thresholds involve trade-offs. A suitable threshold for a low-impact investigation queue may be unsafe for account suspension or network isolation. Record performance by environment, event source and risk class rather than relying only on a single aggregate figure.

    Use representative, time-separated data where possible. Randomly mixing near-duplicate events across training and evaluation sets can exaggerate performance. A later-period holdout is useful because attacker behaviour, infrastructure and normal business activity change over time.

    Treat labels and telemetry as security-critical inputs

    A detector inherits the limitations of its data. Ask:

    • Who defined the ground truth, and how were disagreements resolved?
    • Are incident labels based on completed investigations or only earlier alerts?
    • Does the data include relevant seasons, offices, cloud services and user populations?
    • Which identities, hosts or event sources are missing?
    • Can an attacker influence logs, text, URLs or other model inputs?
    • Does the integration expose secrets, personal information or privileged investigation data?

    Generative systems can be influenced by untrusted content embedded in logs, tickets, webpages or documents. An apparent instruction inside an artefact is data to investigate, not authority to run a command. ASD recommends constrained integrations, appropriate isolation and human oversight for higher-impact actions (ASD — Opportunities for AI in cyber defence).

    Design for drift and adversarial behaviour

    Production performance will change. Software updates alter event formats; a new office changes normal login patterns; attackers adapt to visible controls; and a vendor may update a hosted model without reproducing the original evaluation.

    Monitor:

    • input schema and missing-field rates;
    • alert volume and score distribution;
    • precision and recall on reviewed samples;
    • performance by data source and business unit;
    • overrides, rejected recommendations and response reversals;
    • changes to model, prompt, rules, dependencies and provider terms; and
    • security incidents involving the AI system itself.

    NIST describes evasion, poisoning, privacy and misuse risks across AI system lifecycles in its adversarial-machine-learning taxonomy (NIST AI 100-2e2025). MITRE ATLAS catalogues observed techniques against AI-enabled systems and can help structure threat modelling; it is not a certification checklist (MITRE ATLAS).

    Keep automation bounded and reversible

    Begin in observe-only mode. Let the system recommend or enrich while humans compare results with the established process. Progressively automate only when evidence supports it.

    Safer early actions often have all of these properties:

    • limited effect and short duration;
    • a clear owner and audit trail;
    • an independent check before high-impact execution;
    • a tested rollback path;
    • rate and blast-radius limits; and
    • continued operation if the model or provider is unavailable.

    For example, adding a temporary investigation tag is easier to reverse than deleting data or disabling a workforce account. Isolation, credential revocation, firewall changes and external notifications normally require stronger evidence, explicit authority and a human decision.

    Questions to put to a vendor

    Request evidence that matches the intended environment:

    1. What exact task is the model performing, and what remains rule-based or human-operated?
    2. Which data is collected, retained, transferred or used to improve a provider service?
    3. Can customer data, prompts and outputs be excluded from model training?
    4. How are tenants separated, administrators controlled and access logged?
    5. How was performance measured, on what prevalence and against which baseline?
    6. Can results be broken down by source, environment and error type?
    7. How are model, rule and prompt changes communicated and rolled back?
    8. What happens during service degradation, a provider breach or contract termination?
    9. Can the customer export alerts, evidence, configuration and audit history?
    10. What independent security assessment applies to the actual service being purchased?

    A benchmark percentage without the dataset, label definition, threshold, base rate and operating context is not enough to support a deployment decision.

    A controlled pilot gate

    Before production use, agree on:

    • a named owner and decision authority;
    • the bounded task and prohibited actions;
    • privacy, retention and cross-border-data review;
    • a representative evaluation set and baseline;
    • error and workload thresholds;
    • human review and escalation paths;
    • logging, monitoring and model-change controls;
    • rollback and provider-outage procedures; and
    • a date for reassessment.

    The NIST AI Risk Management Framework organises AI risk work around Govern, Map, Measure and Manage. It is voluntary guidance rather than a guarantee or one-size-fits-all compliance regime (NIST AI RMF Core).

    Where Ozlin can help

    Ozlin can help a small organisation define a bounded AI-assisted security workflow, map data and integrations, establish a baseline, design a pilot and document human review and rollback. Any engagement must define scope, data handling and decision ownership before testing begins. Ozlin does not promise zero-day detection, automatic accuracy improvement, breach prevention or autonomous incident resolution.

    See Cybersecurity services or contact Ozlin to discuss a scoped assessment.

    Related reading: AI chatbots for Australian SMEs.

    This article provides general technical information, not legal, compliance or security assurance. Results depend on data, configuration, people, threat conditions and the specific service evaluated.

    Limitations: Evaluation results do not generalise across models, data, attacks or operations. Measure false positives and negatives, drift, privacy and human workload on representative data; no detection or assurance guarantee is made.

    AI-assistance disclosure

    AI tools assisted with source discovery, outlining and copyediting. A human reviewer must verify every factual statement, source, service claim and publication decision before release. No model, product or control is endorsed by inclusion.

    Primary sources checked

    Source access date: 29 August 2026.

    Article map for AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype, covering Start with a bounded use case, Establish a baseline before adding a model, Measure errors in operational terms and related…
    Article map: Start with a bounded use case; Establish a baseline before adding a model; Measure errors in operational terms; Treat labels and telemetry as security-critical inputs.
    Decision path for AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype, covering Measure errors in operational terms, Treat labels and telemetry as security-critical inputs, Design for drift and advers…
    Decision path: Measure errors in operational terms; Treat labels and telemetry as security-critical inputs; Design for drift and adversarial behaviour; Keep automation bounded and reversible.
    Control and evidence map for AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype, covering Design for drift and adversarial behaviour, Keep automation bounded and reversible, Questions to put to a ven…
    Control and evidence map: Design for drift and adversarial behaviour; Keep automation bounded and reversible; Questions to put to a vendor; A controlled pilot gate.
    Practical checklist for AI in Cyber Defence: How to Evaluate Threat Detection Without the Hype, covering A controlled pilot gate, Where Ozlin can help, AI-assistance disclosure and related review points.
    Practical checklist: A controlled pilot gate; Where Ozlin can help; AI-assistance disclosure; Primary sources checked.
  • Neural Networks Explained: Gradients, Optimisers and Evaluation

    Neural Networks Explained: Gradients, Optimisers and Evaluation

    A neural network is a parameterised function. It transforms input features through layers of weighted operations and nonlinear activations, then produces an output such as a probability, numeric estimate, embedding or generated sequence. “Deep” learning usually means that the model contains multiple representation-building layers; it does not mean that the result is automatically intelligent, correct or suitable for production.

    The useful question is not whether a model sounds advanced. It is whether its inputs, target, evaluation and operating controls fit a defined decision or workflow.

    Article map for Neural Networks Explained: Gradients, Optimisers and Evaluation, covering Follow one training step precisely, Layers learn representations, not explanations, Split data before learning from it and relate…
    Article map: Follow one training step precisely; Layers learn representations, not explanations; Split data before learning from it; Make experiments reproducible enough to investigate.

    Follow one training step precisely

    Training is easier to reason about when four operations are kept separate:

    1. Forward pass: the current parameters transform a batch of inputs into predictions.
    2. Loss calculation: a chosen loss function measures disagreement between predictions and the training target.
    3. Backpropagation or automatic differentiation: the system applies the chain rule through the recorded computation to calculate gradients of the loss with respect to parameters.
    4. Optimiser step: an algorithm such as stochastic gradient descent or Adam uses those gradients and its own state to update the parameters.

    Backpropagation does not select the learning rate and does not itself update weights. Conversely, an optimiser needs gradients but is not the mechanism that differentiates the model. PyTorch documents these responsibilities separately in its autograd mechanics and optimisation API.

    A simplified iteration looks like this:

    clear old gradients
    predictions = model(batch.features)
    loss = loss_function(predictions, batch.targets)
    compute gradients of loss
    optimiser updates parameters
    record metrics and diagnostics

    Frameworks may combine or reorder implementation details, use gradient accumulation, mixed precision or distributed execution, but the conceptual distinction remains important when debugging exploding gradients, stale accumulation or an unexpected update.

    Layers learn representations, not explanations

    A dense layer combines its inputs and parameters, while an activation such as ReLU introduces nonlinearity. Convolutional layers encode useful locality assumptions for grid-like data. Recurrent designs maintain a sequential state. Attention allows elements to condition on other elements and underpins modern transformer architectures.

    These are design biases, not proof that a model has learned the intended concept. A classifier may exploit a watermark, scanner type, postcode proxy or annotation habit instead of the phenomenon the team meant to model. Inspect slices, counterexamples and failure modes rather than treating a high aggregate score as an explanation.

    Start with the simplest credible baseline. A linear model, ruleset or manual workflow can expose whether a neural network creates enough additional value to justify its data, latency, operational and governance costs. Google’s current Machine Learning Crash Course places neural networks alongside data quality, generalisation, production systems and fairness rather than presenting architecture alone as the project.

    Split data before learning from it

    Use separate training, validation and test data, with the split designed around how the model will encounter the real world. Random row splits can leak information when records from the same customer, document, device or time period appear on both sides. For a future-facing forecast, use a temporal holdout. For a system expected to generalise to new organisations, consider holding out entire organisations.

    Any learned preprocessing—normalisation statistics, feature selection, vocabulary building, imputation or synthetic sampling—must be fitted using training data only. Duplicate and near-duplicate examples require deliberate handling. A test set repeatedly consulted during tuning is no longer an untouched final test.

    Choose metrics from the cost of errors. Accuracy can conceal failure on a rare but important class. Depending on the task, review precision, recall, false-positive and false-negative rates, calibration, ranking measures, latency and abstention behaviour. Report uncertainty and results for meaningful cohorts. Do not optimise one convenient metric while leaving the business decision undefined.

    Decision path for Neural Networks Explained: Gradients, Optimisers and Evaluation, covering Layers learn representations, not explanations, Split data before learning from it, Make experiments reproducible enough to inv…
    Decision path: Layers learn representations, not explanations; Split data before learning from it; Make experiments reproducible enough to investigate; Evaluate the system, not only the checkpoint.

    Make experiments reproducible enough to investigate

    Record the code revision, framework and library versions, data snapshot and query, preprocessing configuration, random seeds, model configuration, hardware, checkpoint and evaluation script. Reproducibility is not equivalent to typing one seed. Parallel execution and some hardware kernels can be nondeterministic, while deterministic alternatives can be slower.

    PyTorch explicitly warns that completely reproducible results are not guaranteed across releases, commits, platforms or CPU and GPU execution in its reproducibility notes. TensorFlow likewise documents the software, hardware, input-pipeline and random-state conditions around deterministic operations. Treat those settings as debugging and assurance tools, then measure their performance impact.

    Evaluate the system, not only the checkpoint

    A production model sits inside a larger system: collection, validation, feature computation, inference, human review, logging, fallback, appeal and retraining. Test malformed and missing inputs, out-of-distribution cases, dependency failures, timeouts, rollback and version compatibility. Define what the system must do when confidence is low or a required input is unavailable.

    Monitoring should cover input quality, operational health, outcome quality where ground truth eventually arrives, and effects on people. Drift is not automatically harmful, and a stable input distribution does not guarantee stable outcomes. Establish a named owner, review cadence, incident path and retirement condition.

    The voluntary NIST AI Risk Management Framework organises risk work around Govern, Map, Measure and Manage. Its emphasis on validity, reliability, transparency, privacy and harmful-bias management is a useful reminder that a model can be technically functional yet unsuitable in context.

    Choose tools after defining the constraints

    TensorFlow, PyTorch and other frameworks can all support serious work. Select on required operators, deployment target, team competence, maintenance horizon, observability and ecosystem compatibility—not an outdated claim that one framework is only for research and another only for production. Prototype the riskiest deployment path early and pin versions before a reproducible evaluation.

    Before release, require evidence for:

    • a documented task, user and unacceptable outcome;
    • a leakage-resistant data split and representative test set;
    • a baseline and decision-relevant metrics;
    • cohort and edge-case evaluation;
    • reproducible artefacts and dependency records;
    • privacy, security and access controls;
    • human review or safe fallback where consequences justify it; and
    • monitoring, rollback and accountable ownership.

    For help framing a responsible prototype or evaluation plan, see Ozlin Info’s AI and automation services or contact Ozlin Info.

    Related reading: Machine-learning paradigms and evaluation and reading an older PyTorch HMER project responsibly.


    Control and evidence map for Neural Networks Explained: Gradients, Optimisers and Evaluation, covering Make experiments reproducible enough to investigate, Evaluate the system, not only the checkpoint, Choose tools afte…
    Control and evidence map: Make experiments reproducible enough to investigate; Evaluate the system, not only the checkpoint; Choose tools after defining the constraints; General-information disclaimer.

    General-information disclaimer

    This article provides general technical information, not a guarantee of model accuracy, fairness, safety, regulatory compliance or business outcomes. Independent domain, privacy, legal, security and statistical review may be required for the actual use case.

    AI-assistance disclosure

    AI tools assisted with source discovery, outlining and copyediting. A human reviewer must verify the technical claims, examples, links and risk controls against the intended dataset, framework version and deployment context before publication or use.

    Practical checklist for Neural Networks Explained: Gradients, Optimisers and Evaluation, covering Choose tools after defining the constraints, General-information disclaimer, AI-assistance disclosure and related review…
    Practical checklist: Choose tools after defining the constraints; General-information disclaimer; AI-assistance disclosure; Primary sources checked.

    Primary sources checked

    Source access date: 29 August 2026.

  • Supervised, Unsupervised and Other Machine-Learning Paradigms

    Supervised, Unsupervised and Other Machine-Learning Paradigms

    “Supervised or unsupervised?” is a useful opening question, but it is not a complete project brief. The right learning setup depends on the decision to support, what feedback exists, when it arrives, what mistakes cost and how success can be evaluated on data the system did not learn from.

    Algorithms are also not permanently owned by one paradigm. The same neural architecture can be trained with labelled targets, a self-supervised objective or reinforcement feedback. Recommendation and anomaly-detection systems often combine several approaches. Start with the source of the learning signal rather than a list of fashionable model names.

    Article map for Supervised, Unsupervised and Other Machine-Learning Paradigms, covering Supervised learning uses explicit targets, Unsupervised learning finds structure without target labels, Several important setups si…
    Article map: Supervised learning uses explicit targets; Unsupervised learning finds structure without target labels; Several important setups sit between or beyond the pair; Do not force applications into one bucket.

    Supervised learning uses explicit targets

    Supervised learning trains on examples containing input features and a target label or value. Classification predicts a category or probability; regression predicts a numeric quantity. Examples include identifying an invoice type, estimating delivery time or predicting whether a reviewed transaction belongs to a defined class.

    The target must represent the real decision. Historical labels can contain inconsistent human judgement, policy changes or outcomes produced by the old process. A model may faithfully reproduce those artefacts. Document who created each label, under which rules, and how disagreement and uncertainty are represented.

    Evaluation uses held-out labelled data and task-appropriate metrics. Accuracy is insufficient when classes are imbalanced or errors have asymmetric consequences. Depending on the decision, examine precision, recall, calibration, cost-weighted error, ranking quality and performance by meaningful cohort. Google’s introduction to supervised learning emphasises labelled examples, unseen data and generalisation.

    Unsupervised learning finds structure without target labels

    Unsupervised methods usually operate on unlabelled examples to identify patterns such as groups, lower-dimensional representations or unusual observations. Clustering can support exploration or segmentation, but a cluster is not automatically a real customer type. Results depend on representation, distance, scaling, algorithm and hyperparameters.

    Validate whether a discovered structure is stable and useful outside the training sample. Compare multiple seeds and plausible preprocessing choices, inspect examples with domain experts, and test whether the segmentation improves a downstream decision. An internal cohesion score cannot establish that the groups are fair, causal or commercially meaningful.

    Dimensionality-reduction visualisations require similar restraint. t-SNE was introduced as a method for visualising high-dimensional data; it is not a clustering algorithm, and apparent gaps in a two-dimensional plot are not proof of natural classes. The original t-SNE paper explains its local similarity objective and limitations.

    Several important setups sit between or beyond the pair

    Semi-supervised learning combines a smaller labelled set with a larger unlabelled set. It can reduce labelling demand when the unlabelled data resembles the intended operating distribution, but poor pseudo-labels or a distribution mismatch can reinforce errors.

    Self-supervised learning creates a training signal from the data itself—for example, predicting masked content or contrasting related views—then adapts the learned representation to a downstream task. It still needs careful downstream evaluation; a useful pretraining objective does not guarantee appropriate behaviour in the final context.

    Reinforcement learning learns a policy through interaction, observations, actions and reward. It is not simply supervised learning with delayed labels. Reward design, exploration, environment fidelity, safety constraints and off-policy evaluation can dominate the project. Sutton and Barto’s Reinforcement Learning: An Introduction provides the primary textbook treatment.

    Active learning asks which examples should be labelled next. It can focus limited expert time but must account for sampling bias and the true cost of obtaining a reliable label.

    The current Google machine-learning glossary distinguishes labelled, unlabelled, semi-supervised and unsupervised examples. Use these terms to describe the training signal, not to imply a quality ranking.

    Decision path for Supervised, Unsupervised and Other Machine-Learning Paradigms, covering Unsupervised learning finds structure without target labels, Several important setups sit between or beyond the pair, Do not forc…
    Decision path: Unsupervised learning finds structure without target labels; Several important setups sit between or beyond the pair; Do not force applications into one bucket; Match evaluation to the learning signal.

    Do not force applications into one bucket

    An autoencoder learns to reconstruct or otherwise represent its input. It may contribute an anomaly score, but the score still needs a threshold, representative validation cases and an operational response. Reconstruction error alone does not prove fraud, intrusion or equipment failure.

    A recommender might use supervised ranking from observed outcomes, self-supervised representations, collaborative signals, content features, contextual bandits or business rules. The important questions are which feedback is observed, which is missing, how exposure biases the data, and whether the evaluation captures user and business effects.

    Likewise, “anomaly” can mean a rare statistical point, a rule violation or a high-cost event. A rare point may be legitimate; a harmful event may look common in the available features. Define the review action and tolerated alert burden before selecting an outlier method.

    Match evaluation to the learning signal

    Learning setup Typical evidence Evaluation warning
    Supervised held-out labelled outcomes labels may leak, drift or encode the old policy
    Unsupervised stability, domain review, downstream utility internal cluster scores do not prove real-world meaning
    Semi-supervised labelled holdout plus ablation against labelled-only baseline pseudo-labels can amplify early mistakes
    Self-supervised downstream task performance and transfer tests pretraining loss is not the business metric
    Reinforcement learning policy value, safety limits and online or simulator evidence an exploitable reward can produce the wrong behaviour

    Create train, validation and final test boundaries before feature engineering. Group related records so the same customer, device, document family or future information cannot appear across the boundary. Compare against a simple baseline and include the human or rules-based process where relevant.

    For consequential uses, evaluation also needs privacy, security, fairness, transparency and human-oversight criteria. The voluntary NIST AI Risk Management Framework calls for business context to be mapped, methods and metrics to be documented, and systems to be tested before deployment and monitored afterwards.

    A practical selection sequence

    1. Write the decision, user and unacceptable outcome.
    2. Define the unit of prediction or analysis and when the output is needed.
    3. Inventory available features, labels, feedback delays and collection rights.
    4. Choose the simplest baseline that can be evaluated honestly.
    5. Design a leakage-resistant split and decision-relevant metrics.
    6. Prototype the data and review workflow before scaling the model.
    7. Record limits, owners, monitoring and a safe fallback.

    For help framing an ML pilot or evidence plan, see Ozlin Info’s AI and automation services or contact Ozlin Info.

    Related reading: Neural networks, gradients and evaluation and a transparent AI document-processing ROI example.


    Control and evidence map for Supervised, Unsupervised and Other Machine-Learning Paradigms, covering Do not force applications into one bucket, Match evaluation to the learning signal, A practical selection sequence and…
    Control and evidence map: Do not force applications into one bucket; Match evaluation to the learning signal; A practical selection sequence; General-information disclaimer.

    General-information disclaimer

    This article provides general technical information, not a guarantee of model accuracy, fairness, safety, regulatory compliance or return on investment. Validate the learning setup and evidence requirements for the actual data, decision and affected people.

    AI-assistance disclosure

    AI tools assisted with source discovery, outlining and copyediting. A human reviewer must verify the terminology, links, evaluation design and risk controls against the intended use case before publication or use.

    Practical checklist for Supervised, Unsupervised and Other Machine-Learning Paradigms, covering A practical selection sequence, General-information disclaimer, AI-assistance disclosure and related review points.
    Practical checklist: A practical selection sequence; General-information disclaimer; AI-assistance disclosure; Primary sources checked.

    Primary sources checked

    Source access date: 29 August 2026.