Skip to main content

Diagnostic evaluation

Published Updated 6 min readBy FaultPilot Technologies Pty Ltd

How to evaluate AI-assisted fault diagnosis for mining maintenance

Test the evidence boundary behind generated maintenance guidance before treating fluency as diagnostic confidence.

An AI-assisted maintenance workflow separating authorised sources, generated guidance and qualified human verification

AI-assisted fault diagnosis should be evaluated as an advisory information workflow, not as an autonomous maintenance authority. A fluent answer can still be incomplete, unsupported or outside the controls that apply to the machine and site.

The buyer's task is to inspect what information shaped the output, how limitations are shown and where the qualified maintainer, supervisor and site procedure remain controlling.

Start with the decision the guidance is meant to support

Define whether the product is helping frame the symptom, organise possible causes, retrieve relevant information, build an ordered check plan or preserve a diagnostic trail. Those are different jobs with different evidence requirements.

Avoid accepting the phrase “AI diagnosis” as the scope. Ask which user makes the next decision and what the generated output is allowed to influence.

Inspect the inputs before reading the answer

Record the asset context, symptom, operating conditions, prior work, current findings and enabled technical information supplied to the system. Missing or incorrect input should remain visible because it limits the value of any later guidance.

Test ambiguous, sparse and conflicting reports as well as a clean demonstration. The degraded cases reveal whether the product asks for more information, lowers its confidence or invents a complete story.

Confirm which sources are authorised

Ask how customer documentation enters the system, which organisation can access it, how versions are handled and how an outdated or withdrawn source is treated. Public internet information, organisation-controlled documents and earlier private maintenance records do not have the same authority.

A source label should help the maintainer understand coverage. It does not turn generated guidance into an approved machine procedure.

Require visible source and fallback states

The product should distinguish guidance supported by retrieved authorised information from a reduced-source response or a template that has no governed retrieval behind it. The user needs that state before deciding how much weight to place on the output.

Review the source-state guide for a practical distinction between grounded, reduced-source and template operation.

Keep evidence, hypotheses and instructions separate

A useful workflow identifies observed evidence, possible causes and proposed checks as different objects. It should not restate a hypothesis as a finding or present a generated check as site-approved work.

Ask how each result changes the current hypotheses and how the system records uncertainty, causes ruled out and escalation triggers.

Test the safety and authority boundary

Site procedures, permits, isolations, authorised supervision, competent judgement and current machine documentation remain controlling. Test what happens when a required control is missing, the reported scope exceeds the user's authority or the output conflicts with an approved source.

The safe behaviour may be to stop, request clarification or escalate. It should not be to generate more confident language.

Review data and organisation boundaries

Map the data sent to each model, retrieval, storage and observability provider. Confirm organisation isolation, user access, retention, deletion, incident handling and whether customer information is used for provider or product training.

Require the answer to match the exact product configuration being evaluated. A general provider statement is not evidence of the application's complete data flow.

Measure the workflow, not a polished answer

Use representative synthetic or approved test cases with known evidence boundaries. Review whether the system preserves the symptom, requests missing context, retrieves relevant material, marks source state, proposes bounded checks and leaves a handover-ready trail.

Record failures and limitations as part of the pilot result. Do not score only whether the final answer sounds plausible.

AI-assisted fault diagnosis buyer checklist

  • What exact maintenance decision does the guidance support?
  • Which input fields and prior records shape the output?
  • Which sources are authorised, versioned and organisation-scoped?
  • Can the user see grounded, reduced-source and template states?
  • Are observations, hypotheses, checks and findings kept distinct?
  • What stops or escalates work when controls or evidence are missing?
  • Which providers process the data, for what purpose and for how long?
  • How are incorrect, conflicting and low-evidence outputs recorded?
  • Can the resulting work be reviewed and handed over to another qualified person?

Compare this checklist with the systematic fault finding process and FaultPilot's bounded diagnostic intelligence model.

Review the diagnostic evidence boundary →