We don’t know how much it cost to create, what potential cybersecurity risks it posed or even whether it saved money.
The Department of Health, Disability and Ageing created an AI-powered fraud detection tool that was only correct 22% of the time, cost an unknown amount of money to produce and turned up so many false positives that it became a drain on resources. It was used for more than a year.
Yesterday the Australian National Audit Office published its performance audit of artificial intelligence and Medicare benefits integrity.
It looked at both the adoption of AI for Medicare compliance purposes and at whether the DoHDA was considering the role that AI might play in driving incorrect or fraudulent medical billing practices.
The bulk of the report examined an AI-enabled model which DoHDA used in its Medicare fraud detection system between July 2024 and December 2025.
The problems with the model started from its inception
One of the biggest questions raised by the audit was how much the intervention – which was implemented alongside six other, non-AI powered models – cost.
DoHDA, at least, does not appear to have kept track.
“The 2023-24 federal Budget measure Strengthening Medicare — improving Medicare integrity included costings [of $29.8 million over four years] at a high level but did not specify costs for developing the fraud detection system,” the ANAO wrote.
“DoHDA advised the ANAO in April 2026 that the compliance cost model had not been updated to include relevant data that would have allowed for its use in this instance due to structural changes and a review of the broader Benefits Integrity Division operating model.
“DoHDA did not otherwise develop a budget for its work on the fraud detection system or calculate the actual cost of its development or use.”
The AI-powered tool also predated DoHDA’s enterprise AI governance arrangements. This meant it did not get the scrutiny which it perhaps warranted.
“In November 2024 a use case assessment was completed for the system as part of a whole‑of‑government AI pilot before the introduction of the Policy for the responsible use of AI in government,” the audit reads.
“The assessment was not recorded in DoHDA’s AI use case register.
“Although the system used personal information, the system was assessed as having low inherent ethical risk.
“As a result, privacy and legal risks were not subject to more detailed assessment, and legal advice was not sought.”
Related
A privacy threshold assessment conducted in May 2026, though, identified that there were high-risk factors in the tool, related to using and disclosing personal information to profile or predict the behaviours of individuals.
Again, because the implementation of the tool predated the AI governance rules, it was not managed within DoHDA’s enterprise governance and security framework.
As in: the teams responsible for cyber security were unaware of the existence of the AI fraud detection tool.
A cyber security risk assessment was never completed.
At the time of the audit’s writing, DoHDA was seeking advice from the Australian Government Solicitor on the question of privacy threshold risk.
Then there was the model itself
In short, the AI-powered model was not great at its job.
It was very good at flagging potential cases for review, but it was only correct about 22% of the time, putting it firmly in the middle of the pack when compared to the six other, non-AI models it was implemented alongside.

“Eight potential non-compliance or fraud matters were created and referred for preliminary analysis using, in part, the AI-enabled model in the MBS fraud detection system,” the ANAO wrote.
“While ad hoc preliminary performance analysis showed the AI-enabled model was efficacious, it was not used after December 2025, when it was determined that resourcing required for manual assessment of fraud and non-compliance ‘signals’ could not keep up with the volume of potential matters generated.
“There was no corresponding assessment of the benefits and costs of this decision.”
It was also unclear whether DoHDA considered how inherent biases in datasets may have influenced the behaviour of the model; whether, for instance, it would unfairly discriminate against certain individuals, communities or groups.
“DoHDA advised the ANAO in March 2026 that consideration of model behaviour occurred but was largely undocumented,” the ANAO wrote.
“DoHDA also advised it had drafted a monitoring framework that proposes standardised metrics to alert for over-representation of certain groups in results, including required actions when this occurs, but that these improvements will not be implemented for the fraud detection system until software is approved.”
What’s more, evidence of verification and validation was not consistently captured; this earned DoHDA its single red “not implemented” mark in the ANAO audit.
“DoHDA advised that business approval prior to deployment, and for changes, was not consistently captured,” the ANAO wrote.
“In some cases, approval was provided verbally only.
“The February 2026 Requirements and Change Management Framework and a July 2026 Governance Framework and Agreement establish requirements and expected practice for capturing business approval within the change and release workflow.”
What it did do
Because the specific costs were not documented, it’s difficult to say whether the AI-enabled fraud detection model was cost-effective.
What the ANAO did establish, though, was that the value of the fraud and non-compliance associated with the eight matters that the AI model likely helped detect was around $5.2 million. However, even this came with a caveat.
“Analysis identified that the eight matters combined results from different models, including non-AI models … and results reporting did not differentiate,” the audit read.
“The analytic approach therefore did not generate useful insights into the unique benefits of AI usage in the system.”
Ultimately, the AI model was deemed a “drain on the resources” of the department due to high volume of potential fraud or serious non-compliance that it flagged for manual review.
So DoHDA simply stopped using it.
“DoHDA advised the ANAO in February 2026 that in December 2025 a decision was made to cease using the AI-enabled model in favour of a logistic regression model because the logistic regression model also performed effectively, and had the benefit of using additional parameters and a threshold to enable direct control over the volume of selections for manual assessment,” the ANAO wrote.
“DoHDA did not assess benefits and costs of the decision nor document the change from the AI-enabled model to the logistic regression model.
“DoHDA advised the ANAO in May 2026 that a business decision was subsequently made to use the logistic regression model based on the model’s comparable efficacy and the available resourcing to perform manual assessment.
“DoHDA did not compare efficacy of the models nor assess model output against available resourcing.”
Read the full audit here.



