What an AI agent can do in Medicare billing that a billing team can't
More than half of the 6,046 items in today's Medicare Benefits Schedule have changed since mid-2021. What an AI agent can do in MBS billing that people can't, the numbers behind it, and where it has to stop.
Mitch Flindell · 28 September 2026 · 21 min read · Journal
People keep asking us what AI can do in a regulated back office that a good team can't. It's a fair question. The honest answer is narrower than the marketing, and more useful.
An agent isn't smarter than an experienced billing officer. It differs in ways that come from how it works rather than how clever it is, and some of those differences can be measured. So we measured what we could in one place: Medicare Benefits Schedule (MBS) billing, the daily work of practice staff, specialists' rooms, day hospitals and medical billing bureaus.
We downloaded all 32 MBS XML data releases from July 2021 to August 2026 and counted what changed3. We read the Australian National Audit Office's audit of AI and Medicare integrity, published on 21 September 20261, and the 2023 Independent Review of Medicare Integrity and Compliance, known as the Philip Review2. Then we looked for the best evidence on how consistently people apply rules. Where a number is our own calculation, we say so, show the method and publish the data at the end of this page.
Key numbers
- ▪6,046 items are in the Medicare Benefits Schedule XML file for 1 August 2026 (Pragmatic AI count)3.
- ▪2,666 of those items (44%) were added or had their descriptor reworded after July 2023, and 3,389 (56%) after July 2021 (Pragmatic AI analysis of 32 MBS XML releases)3.
- ▪The MBS Book operating from 1 July 2026 runs to 1,760 pages with about 770 explanatory notes4. At the published average adult reading speed for non-fiction, 238 words per minute, reading its roughly 694,000 words once would take about 48.6 hours (Pragmatic AI calculation)413.
- ▪482.5 million Medicare services were paid in 2025-26, worth $35.0 billion in benefits7.
- ▪$1.5 billion to $3 billion a year is the Philip Review's indicative estimate of MBS provider non-compliance on a conservative definition2, repeated by the ANAO in 20261. It is an estimate, not a measured loss.
- ▪3,779 cases of MBS fraud and non-compliance were identified in 2024-25, with $44.7 million in debt raised1.
- ▪27% of the 4,820 potential matters flagged by the department's detection models between August 2023 and August 2025 were reviewed (Pragmatic AI sum of ANAO Table 5.1)1.
- ▪At least 25% of practitioners and practice managers in a 2013 Department of Human Services survey did not check whether the item numbers they chose were the ones billed (786 respondents, reported in the department's 2017 billing assurance toolkit)8.
Why Medicare billing is the fairest test
We considered four Australian environments our clients work in: day surgery and private hospital billing, MBS billing in practices and bureaus, NDIS provider claiming, and workers compensation claims. The NDIS has strong audit data of its own. MBS billing won because, of the four, it is the one where the item schedule itself is published as structured data with every release3. That means claims about how big the schedule is, and how fast it moves, can be counted instead of asserted.
The XML is not the whole rulebook. It holds item descriptors, fees and dates. The explanatory notes sit in the MBS Book4, the legal force sits in the legislation, and where they differ the legislation prevails. Any serious check has to use all three. The XML is simply the part that can be counted release by release.
MBS billing is also where accountability is written down most clearly, which matters for the second half of this piece. And it is large. According to the department's Medicare annual statistics, 482.5 million services were paid in 2025-267. By our subtraction of 440.0 million out-of-hospital services from the total, about 42.5 million were in-hospital services, across all hospital settings7. That is a count of services, not of the accounts any one kind of provider handles.
The cost of the status quo is real, and smaller than the headlines
The Philip Review concluded that, on a conservative definition, it is "entirely feasible" that MBS provider non-compliance is worth $1.5 billion to $3 billion a year2. It was careful to call its analysis partial and based on limited data, not definitive. The ANAO repeated the range in its September 2026 audit1. It is the range we use, as an indication of scale.
You will also see $8 billion or $10 billion quoted. Those figures come from Dr Margaret Faux's bottom-up work. The Philip Review took them seriously but with caveats: in the review's assessment they rest on billing data from a small and biased sample, and on a far broader definition of non-compliance than the department uses2. The review judged that about $3 billion of that estimate appeared to share the department's definition2. The larger numbers are largely an argument about definitions, not a measured loss.
What is recovered is much smaller than any of these estimates, and it is measured differently. In 2024-25 the department raised $44.7 million in debt against 3,779 identified cases of MBS fraud and non-compliance1.

Not all of the error runs one way. The Philip Review noted that underbilling "was indicated to me as a growing feature of the system"2. Practitioners who fear an audit may leave legitimate items unbilled. The review did not put a value on underbilling, so its $1.5 billion to $3 billion range should not be read as a measure of billing error in both directions.
The people doing this work are stretched. In the RACGP's Health of the Nation 2025 survey, 77% of GPs were dissatisfied with the amount of administration in their work, up from 70% a year earlier11. Jobs and Skills Australia's 2025 Occupation Shortage List rates clinical coders as in shortage nationally and in every state and territory, and medical receptionists and health practice managers as not in shortage12. Those national ratings don't describe every practice, but they suggest the pressure is less about finding front-office staff than about how much reading and checking the job demands.
An agent can check every account, not a sample
On the regulator's side, checking is rationed. The Philip Review found that, owing to capacity constraints, less than 20% of identified risks progressed to risk assessment, and fewer than 10% of risks brought to regulators reached any kind of treatment pathway2. It counted only 40 to 60 compliance projects a year examining cohorts of providers, in a system with 176,000 practitioners who claim2.
The ANAO's 2026 audit shows the same limit inside the department's own analytics. Its Table 5.1 lists 4,820 potential matters identified by seven detection models between August 2023 and August 2025. By our sum of the ANAO's figures, 1,319 were reviewed, or 27%1. Potential matters can include repeat signals for the same provider, and models 6 and 7 are versions of one model used in different periods1.
| Model | Flagged | Reviewed | Review rate | True positive | False positive | True positive rate |
|---|---|---|---|---|---|---|
| Model 1 (AI-enabled) | 597 | 82 | 14% | 18 | 64 | 22% |
| Model 2 | 990 | 119 | 12% | 12 | 107 | 10% |
| Model 3 | 600 | 287 | 48% | 251 | 36 | 87% |
| Model 4 | 675 | 272 | 40% | 16 | 256 | 6% |
| Model 5 | 1,016 | 345 | 34% | 158 | 187 | 46% |
| Model 6 | 598 | 127 | 21% | 63 | 64 | 50% |
| Model 7 | 344 | 87 | 25% | 41 | 46 | 47% |
| Total (our sum) | 4,820 | 1,319 | 27% | 559 | 760 | n/a |
The pressure came before the models. In December 2023 the department counted the providers showing claiming behaviours known to be associated with fraud, and found more than its staff could look at1. That is why it built the AI-enabled model.

Inside practices, the most recent published evidence we found is old. In a 2013 Department of Human Services survey of 786 practitioners and practice managers, reported in the department's 2017 billing assurance toolkit, at least 25% said they did not check whether the item numbers they chose were the ones billed, and only 77% of practitioners were told when their chosen item was changed8. The survey may not describe current practice.
Today: checked as time allows
- ✕Staff check what they can, against the rules they remember.
- ✕Schedule changes arrive as bulletins to be read and remembered.
- ✕Errors surface later, as rejections, letters or audits.
With an agent: every account, on arrival
- ✓Every account checked against the item text and notes in force on its date of service.
- ✓Every check logged with the rule text and version it relied on.
- ✓Exceptions go to a person, with the evidence attached.
An agent changes the arithmetic. Once it is set up, the added cost of checking one more account is small, so there is less reason to sample. That doesn't make every check right. Its coverage, accuracy, running cost and the point where a person steps in all need measuring in the practice where it runs. What it can do is examine every account, and put the mistakes that remain in one place where someone can see them.
An agent can look up the rule text for every account
The August 2026 MBS data file lists 6,046 items. Their descriptors alone run to 276,066 words (our count)3. The MBS Book operating from 1 July 2026 is 1,760 pages, with about 770 explanatory notes and roughly 694,000 words of text by our extraction4.
Pragmatic AI calculation: a 2019 meta-analysis of 190 reading studies put the average adult silent reading speed for non-fiction at 238 words per minute13. At that pace, one pass of the MBS Book takes about 48.6 hours, or six and a half 7.5-hour working days. The item descriptors alone take about 19.3 hours. That is an illustration of volume, not a measure of how long an experienced officer needs to check one claim: nobody reads the whole book to bill a colonoscopy. It does show why nobody holds all of it in their head.
And the text doesn't sit still. The department published 31 updated versions of the schedule data between August 2021 and August 20263. Of the 6,046 items in the August 2026 file, 2,666 (44%) were new or had their descriptor reworded after July 2023, and 3,389 (56%) after July 20213. Our method: normalise the whitespace in every descriptor, compare each consecutive pair of releases, and count an item in the August 2026 file if it first appears, or its descriptor text differs, at any release after the start date. The Philip Review reported being told in 2023 that "some 3,000 Medicare items underwent some change over the last 2-3 years"2. That was a different period and a looser definition, but it points the same way.
| Period | Releases | Added | Removed | Reworded | Items changed | Share since Jul 2021 |
|---|---|---|---|---|---|---|
| Jul 2021 to Jul 2022 | 8 | 171 | 355 | 452 | 865 | 7.8% |
| Jul 2022 to Jul 2023 | 9 | 318 | 350 | 669 | 1,329 | 21.8% |
| Jul 2023 to Jul 2024 | 5 | 204 | 158 | 664 | 1,017 | 31.1% |
| Jul 2024 to Jul 2025 | 4 | 76 | 43 | 1,788 | 1,893 | 52.2% |
| Jul 2025 to Jul 2026 | 4 | 70 | 14 | 436 | 518 | 56.0% |
| Jul 2026 to Aug 2026 | 1 | 1 | 0 | 0 | 1 | 56.1% |
Not every change alters what can be billed. Some are wording. Others are not: on 1 March 2025, over 800 items lost their 85% out-of-hospital benefit and became payable only in hospital6. In the XML, that release reworded 1,105 descriptors, and 829 of them gained the "(H)" marker for hospital-only services (our count)3. It is the steepest step in the chart, and it matters to anyone billing those procedures.
Many rules also depend on other rules. By a keyword count we publish below, 2,062 of the 6,046 item descriptors, about one in three, refer to another item number or use a phrase that limits how often or how soon an item can be claimed, such as "not more than" or "once in any 5 year period"3. A match is a pointer, not a legal reading, and limits that sit in explanatory notes or legislation are not counted. People carry the items they use most in their heads. An agent can be built to retrieve the exact descriptor and notes from the version that applied on the date of service, every time, and quote them.
An agent can apply the same check on the 500th account
We found no published study of how consistently Australian billing staff choose MBS items. The best evidence on human consistency comes from clinicians and coders overseas. Read it as a signal about people, not a measurement of billing.
In a US study of 21,867 visits to 204 primary care clinicians, published in JAMA Internal Medicine in 2014, the adjusted odds of prescribing antibiotics for respiratory infections were 26% higher in the fourth hour of a clinic session than in the first14. A 2019 US study in JAMA Network Open found breast cancer screening was ordered for 63.7% of eligible patients seen at 8 am and 47.8% of those seen at 5 pm15. Both are observational: they show decisions drifting with the clock, not why.
Agreement between experts is lower than most people assume. In a German study of ICD-10 coding published in 2008, coding specialists given the same discharge letters agreed on the full principal diagnosis code in 46.8% of pairs16. The authors blamed the "innumerous coding rules"16. In clinical research, a 2025 meta-analysis of 93 older studies found a pooled error rate of 6.57% for data abstracted by hand from medical records, against 0.14% for double data entry17. Those are different methods in different studies, a signal about manual transcription rather than a comparison with an agent.
| Study | Setting | Sample | Finding |
|---|---|---|---|
| Linder et al. 2014 | US primary care, 23 practices | 21,867 respiratory infection visits, 204 clinicians | Adjusted odds of an antibiotic prescription 1.26 times higher in the fourth hour of a session than the first14 |
| Hsiang et al. 2019 | US primary care, 33 practices | 19,254 patients eligible for breast screening | Screening ordered for 63.7% of patients seen at 8 am, 47.8% at 5 pm15 |
| Stausberg et al. 2008 | German ICD-10 diagnosis coding | 13 coding specialists, 12 discharge letters | Pairs agreed on the full principal diagnosis code 46.8% of the time16 |
| Garza et al. 2025 | Clinical research data, meta-analysis | 93 studies published 1978 to 2008 | Pooled error rate 6.57% for manual abstraction from records, 0.14% for double data entry17 |
| Soroush et al. 2024 | GPT-4 generating US codes (a model, not people) | US code descriptions, no tools | Exact match 33.9% for ICD-10-CM and 49.8% for CPT18 |
An agent built on fixed checks and retrieved rule text can run the same checks at 6 am and 6 pm, on the first account and the five-hundredth. That is a design property, not a given: a generative model does not always give the same answer twice, so the parts that must be repeatable should be deterministic lookups, and the model's output should be tested and logged. And consistency is not correctness. Where a rule is genuinely ambiguous, an agent will be consistently uncertain, and that uncertainty should go to a person rather than be resolved by guesswork.
An agent can check the moment an account arrives
The Philip Review described the current system for risk identification as "predominantly project based rather than continuous and proactive"2, and put "the lack of emphasis in decision support pre-claims and pre-payment" at the core of the problem2. It recommended linking practice software to Medicare's rules so that a likely error gives "live, real-time feedback" to a staff member, who then confirms with the practitioner2. That is before the claim goes, when fixing it is cheapest.

☛ INV-7026 · fictional example
06:52
A colonoscopy account from Tuesday 14 July arrives with the procedure report. Item entered: 32222.
06:53
Checked against the item text in force that day. Report says asymptomatic, positive screening FOBT; since 1 July that is item 32219. Query drafted with both texts.
09:10
The billing officer puts the query to the gastroenterologist, who confirms 32219. Both names go on the account.
INV-7026 is invented, but the rule change in it is real. Until 30 June 2026, a colonoscopy after a positive faecal occult blood test was billed under item 322223. From 1 July, an asymptomatic patient referred after a positive screening test falls under new item 32219, and 32222 no longer lists that indication53. A symptomatic patient's diagnostic colonoscopy can still fall under 32222. That distinction lives in the clinical record, which is why the query goes to the practitioner.

Timing matters for money as well as accuracy. An incorrect Medicare payment must be repaid. The department says an administrative penalty may apply to the debt, "however, no penalty will apply to voluntary acknowledgements"25269. An error caught and corrected before the claim is paid never becomes a debt at all.
Arrival is also when an agent can cross-check history. Item 32223, for example, is "applicable once in any 5 year period"3. An agent can search every prior account in a practice's or bureau's own records in seconds. It cannot see claims made elsewhere; only Services Australia's assessment does that. It narrows the surprises. It doesn't remove them.
An agent can leave a complete record of every step
Medicare puts the obligation on the practitioner. The department's billing assurance toolkit says practitioners are responsible for all claims made under their provider number or name, "even if the practice software was used to facilitate the process, for example, by automatically pre-populating the MBS item numbers"8.

Responsibility can also reach the business around the practitioner. Since 1 July 2019, the Shared Debt Recovery Scheme has allowed a compliance debt, in some cases, to be shared with an employer or other party that has an arrangement with the practitioner1027. That needs grounds that the other party could have influenced the claim or benefited from it, and a split that is fair and reasonable; voluntary acknowledgements and routine corrections are excluded27.
A person's record of a check is usually initials or a tick. An agent's record can hold the source document, the exact rule text and version it relied on, what it found, what it did, and who approved the exception, each with a timestamp. When a question comes months later, the answer is a lookup rather than a search through memory and email.
What an agent cannot do, and must not be allowed to do
It cannot decide what service was provided. The Philip Review is plain: "Legislatively, responsibility lies with healthcare providers to ensure Medicare claiming occurring under their provider number meets all requirements"2. Ahpra's 2024 guidance on AI says the same for clinical work: the practitioner "remains responsible" and "must apply human judgment to any output of AI"19. An agent can put the evidence and the rule side by side. The workflow we build sends uncertain clinical facts and item choices to the practitioner to confirm, and records that decision.

The agent does the reading and the checking. The practitioner stays responsible for what gets billed.
It must not recall rules from memory. In a 2024 NEJM AI study, GPT-4 asked to generate US billing codes from their descriptions matched the correct ICD-10-CM code 33.9% of the time and the correct CPT code 49.8% of the time18. That was not a test of MBS billing, and it tested a model working without tools. The lesson is still the right one: a billing agent should look up the current authoritative text, quote it, show its working, and be tested on the task it will actually do.
Its flags are not findings. In the ANAO's account of the department's preliminary analysis, 18 of 82 reviewed signals from the AI-enabled detection model were classified as true positives, a rate of 22%1. Even those were potential matters for further assessment, and no money had been recovered as of February 20261. The model was used from July 2024 to December 20251. A flag is a reason for a qualified person to look, nothing more.
It must not be built to bill more. The ANAO reports that a draft departmental paper noted "an increase in software packages maximising MBS claiming by using AI to generate MBS item suggestions", and that in March 2026 the department told the auditors it cannot tell which providers have used an AI billing assistant1. The department also said AI was not then assessed as a demonstrated Medicare integrity risk in its own right, though it may increase the scale or speed of existing risks1. An agent worth having checks proposed claims for accuracy and clinical appropriateness whichever way the amount moves, catching underbilling as readily as overbilling, and hands the choice to the practitioner.
The surrounding rules point the same way. The OAIC's October 2024 guidance says privacy obligations apply to personal information put into an AI system and to what it produces, and recommends against entering personal information into publicly available generative AI tools20. From 10 December 2026, an APP entity must say in its privacy policy when a computer program uses personal information to make, or do something substantially and directly related to making, a decision that could reasonably be expected to significantly affect an individual's rights or interests2421. That can cover a program that prepares a decision a person then makes. The TGA says software that only organises and processes Medicare claims may be excluded from medical device regulation under exclusion 14G, but only if every function meets the criteria; anything with a clinical function needs its own assessment22. The National AI Centre's Guidance for AI Adoption sets out six essential practices, including human oversight23. None of that makes an agent unsafe. It means the operator carries the governance, and needs to be able to show it.
What this means for a practice, day hospital or bureau
The split is clean. The reading, the checking against the item text and notes in force, the cross-checking against history and the record-keeping are work an agent can be designed to do on every account as it arrives, with its coverage and accuracy measured where it runs. Item choice where the facts are unclear, ambiguous rules and responsibility for the claim stay with the practitioner and the qualified staff who support them.
That is how we design case engines: the engine prepares the routine accounts for review, sends the rest to the right person with the rule and the evidence attached, and records everything. If you'd like to see what that could look like for your accounts, start with a free conversation about how they arrive today.
Method and data
Everything we counted can be recounted. We downloaded the XML file linked on each of the 32 MBS Online downloads pages from July 2021 to August 2026 (where a page links several re-issued files, the first one listed)3. Each record's key is ItemNum, plus SubItemNum where present. Before comparing, we collapsed all whitespace in the Description field to single spaces and trimmed it; without that step, line-break differences add about a dozen spurious changes. An item in the 1 August 2026 file counts as new or reworded after a start date if, at any later release, it was absent from the previous release or its normalised descriptor differed.
The frequency and cross-reference count applies two case-insensitive patterns to the normalised August 2026 descriptors: any "item" or "items" followed by a two to five digit number, or a limit phrase such as "not more than", "once per", "once in", "in a calendar year", "in the preceding 12" or "per lifetime" (the full patterns are in key-figures.csv). The explanatory-note count is the number of distinct note identifiers (such as TN.8.152) that start a line followed by heading text in the July 2026 MBS Book, extracted with PyMuPDF; the same extraction gives 693,965 words. The Table 5.1 totals are sums of the ANAO's seven rows.
The files are free to reuse under CC BY 4.0 with a link to this page. The MBS data itself is published by the Department of Health, Disability and Ageing, and the MBS Online files are general information, not legal documents; the legislation prevails3.
Sources
- 1Artificial Intelligence and Medicare Benefits Integrity (Auditor-General Report No. 5 of 2026-27) · Australian National Audit Office, 2026
- 2Independent Review of Medicare Integrity and Compliance: Final report (Philip Review) · Department of Health and Aged Care, 2023
- 3MBS Online downloads: MBS XML files, July 2021 to August 2026 · Department of Health, Disability and Ageing, 2026
- 4Medicare Benefits Schedule Book, operating from 1 July 2026 · Department of Health, Disability and Ageing, 2026
- 5MBS Online: July 2026 News, summary of changes · Department of Health, Disability and Ageing, 2026
- 6MBS Online: March 2025 News, change to hospital only services · Department of Health and Aged Care, 2025
- 7Medicare annual statistics: State and territory (2009-10 to 2025-26) · Department of Health, Disability and Ageing, 2026
- 8Medicare billing assurance toolkit: strategies to minimise risk · Department of Health, 2017
- 9How to comply with Medicare obligations · Department of Health, Disability and Ageing, 2026
- 10Managing Health Provider Compliance (Auditor-General Report No. 17 of 2020-21) · Australian National Audit Office, 2020
- 11General Practice: Health of the Nation 2025 · Royal Australian College of General Practitioners, 2025
- 122025 Occupation Shortage List · Jobs and Skills Australia, 2025
- 13How many words do we read per minute? A review and meta-analysis of reading rate · Journal of Memory and Language (Brysbaert), 2019
- 14Time of day and the decision to prescribe antibiotics · JAMA Internal Medicine (Linder et al.), 2014
- 15Association of Primary Care Clinic Appointment Time With Clinician Ordering and Patient Completion of Breast and Colorectal Cancer Screening · JAMA Network Open (Hsiang et al.), 2019
- 16Reliability of diagnoses coding with ICD-10 · International Journal of Medical Informatics (Stausberg et al.), 2008
- 17Error rates of data processing methods in clinical research: a systematic review and meta-analysis · International Journal of Medical Informatics (Garza et al.), 2025
- 18Large Language Models Are Poor Medical Coders: Benchmarking of Medical Code Querying · NEJM AI (Soroush et al.), 2024
- 19Meeting your professional obligations when using Artificial Intelligence in healthcare · Ahpra, 2024
- 20Guidance on privacy and the use of commercially available AI products · Office of the Australian Information Commissioner, 2024
- 21Consultation on Guidance for Transparency in Automated Decision Making · Office of the Australian Information Commissioner, 2026
- 22Understanding the health facility management software exclusion · Therapeutic Goods Administration, 2026
- 23Guidance for AI Adoption: implementation practices · Department of Industry, Science and Resources (National AI Centre), 2025
- 24Australian Privacy Principles guidelines, Chapter 1: APP 1 Open and transparent management of personal information · Office of the Australian Information Commissioner, 2025
- 25Medicare debts and penalties · Department of Health, Disability and Ageing, 2026
- 26Voluntary acknowledgement of incorrect payments · Department of Health, Disability and Ageing, 2026
- 27Shared Debt Recovery Scheme · Department of Health, Disability and Ageing, 2026