#1 in AI Orthopaedic Coding
You audit an AI coder the same way you audit a human coder: with a baseline sample, a defined accuracy standard, an independent reviewer, and a documented remediation loop. Three things are different. You audit before go-live, you audit continuously rather than annually, and you audit the system's confidence calibration, not just its output.
The reason is structural. A human coder makes errors that are random and individual. An AI coder makes errors that are systematic and reproducible, which means a single unaddressed logic gap propagates across every chart it touches. That is the compliance risk. It is also the compliance opportunity, because systematic errors are findable and fixable in a way that human variance never has been. Here is the framework orthopaedic practices should use.
Why "is it accurate?" is the wrong first question
The question practice leaders ask vendors first is almost always some version of "what is your accuracy rate?" It is a reasonable instinct, but it comes up short, for three reasons.
Accuracy is undefined until you name the denominator. Accuracy on what? The primary CPT? Every CPT on the claim? Modifiers? ICD-10 to the full character? A vendor reporting 96% accuracy on primary procedure code selection and a vendor reporting 96% accuracy on complete claim composition are describing very different products.
Accuracy against what standard? Against the practice's historical coding, which may itself be undercoded? Against an independent certified auditor's determination? Against payer adjudication outcomes? These produce different numbers, and the first one, agreement with your existing coders, measures conformity rather than correctness.
Accuracy on whose case mix? Orthopaedics is not one specialty. A system validated on high-volume office E/M and knee arthroscopy will not perform identically on complex spine, revision arthroplasty, or hand microsurgery. A single blended accuracy figure hides exactly the service lines where errors are most expensive.
The better first question: "Show me the audit methodology, and show me the results broken out by service line and by error type."
The compliance context orthopaedic practices are operating in
Two data points frame why this matters now.
First, the federal error picture. CMS reported a Medicare fee-for-service improper payment rate of 6.55%, or $28.83 billion, for FY2025, down from 7.66% and $31.70 billion in FY2024, and the ninth consecutive year below the 10% threshold. The mechanism behind those errors is the part orthopaedic practices should internalize. Per CMS, most improper payments occurred where a reviewer could not determine if a payment was proper because of insufficient documentation. Not fraud. Not aggressive coding. Documentation that did not carry the claim.
Second, the adoption picture. Experian Health's 2025 State of Claims survey found that while 62% of providers are now familiar with AI claims solutions, up from 28% the prior year, only 14% currently use them. Among those who do, 69% report reduced denials or improved resubmission success. The gap between familiarity and adoption is largely a governance gap. Leaders can see the case and cannot yet articulate how they would defend the decision to a payer, an auditor, or a board.
An audit framework is what closes that gap. It is not a compliance tax on adoption. It is the thing that makes adoption defensible.
Phase 1: The pre-deployment baseline audit
Before a single AI-generated code reaches a claim, establish what you are comparing against. Most practices skip this step and then cannot answer the only question that matters at renewal: did this help?
Build a retrospective validation set
Pull a stratified sample of closed, adjudicated encounters, stratified by service line rather than randomly sampled. A workable structure for a mid-sized orthopaedic group:
- Office E/M, with levels 3, 4, and 5 scored separately, this is where downcoding hides
- Fracture care
- Arthroscopy
- Arthroplasty
- Spine
- Hand and foot and ankle
- Injections and ancillary procedures
Have an independent certified coding auditor, not your own coding staff and not the vendor, determine the correct code set for each encounter from the documentation alone. That is your ground truth. It is also, notably, the first time most practices learn their real baseline accuracy, which is frequently a more uncomfortable number than the AI's.
Score three things, separately
- Agreement rate. How often does the AI's output match ground truth exactly?
- Error direction. Of the disagreements, how many are undercoding (revenue lost) versus overcoding (compliance exposure)? These are not equivalent risks and should never be netted against each other into a single accuracy figure.
- Error type. Wrong primary code, missing secondary code, missing modifier, wrong modifier, wrong ICD-10 specificity, missing laterality. Error type is what tells you whether a problem is fixable.
Run the same sample against your current human coding output. Now you have a real comparison rather than a vendor benchmark.
Phase 2: Confidence calibration, the audit step unique to AI
This is the step that separates a serious AI coding audit from a human coding audit, and the one most practices do not know to ask about.
A well-designed autonomous coding system does not just output a code. It outputs a code and a confidence level, then routes low-confidence cases to human review. The value of that design depends entirely on whether the confidence signal is calibrated, meaning that when the system says it is 95% confident, it is right about 95% of the time.
A miscalibrated system is dangerous in a specific way. It is confidently wrong, which means its errors bypass the human review queue that exists to catch them. Overconfidence is worse than low accuracy, because low accuracy with honest confidence still routes the hard cases to a person.
How to audit it
Bucket the validation set by the system's stated confidence, for example 99% and above, 95 to 99%, 90 to 95%, and below 90%. Compute actual accuracy within each bucket. Plot stated confidence against observed accuracy. A calibrated system tracks the diagonal. A system that reports 99% confidence and delivers 91% accuracy in that bucket has a governance problem regardless of its blended accuracy figure.
Then ask the operational question: at what confidence threshold does the system route to a human, and who set that threshold, you or the vendor? That threshold is a business decision about your risk tolerance. It should be yours.
Phase 3: Continuous monitoring, not annual audits
Annual coding audits are a reasonable cadence for human coders because human error is roughly stationary. AI systems are not stationary. They are updated, retrained, and re-tuned. Payer rules change. The code set changes every October 1 and every January 1. A system validated in March may behave differently in November.
A workable monitoring structure:
- Weekly: automated distributional monitoring. Track E/M level distribution, modifier frequency (especially 25, 59 and the X modifiers XE, XP, XS, XU, plus 57, 24, 58, 78, and 79), and code-pair frequency against a rolling baseline. A sudden shift in modifier 25 rate is a signal well before it is a denial.
- Monthly: a small independent sample. Thirty to fifty charts, stratified, reviewed against ground truth. Enough to catch systematic drift, small enough to actually happen.
- Quarterly: a full audit by service line, with error-type breakdown and a documented remediation log.
- Event-triggered: re-audit after any vendor model update, any annual code set change, any payer policy change affecting a high-volume code, and any denial-rate anomaly.
Tie the monitoring to outcomes you already track: clean claim rate, initial denial rate, final denial rate, and days in A/R. If AI coding is working, those move. If they do not move, the accuracy number is not measuring anything that matters.
Phase 4: Document the governance, not just the results
If a payer audit or an OIG inquiry arrives, the question will not be "did you use AI?" It will be "what was your process for ensuring the codes you submitted were supported by the documentation?" That is the same question it has always been. The answer needs to be written down.
A defensible AI coding governance file contains:
- The pre-deployment validation methodology and results, by service line
- The named human accountable for coding compliance, a person, with credentials
- The confidence threshold policy and who approved it
- The human review workflow: what gets reviewed, by whom, and what authority the reviewer has to override
- The monitoring cadence and the actual monitoring records
- The remediation log: every identified systematic error, the date identified, the fix, and the verification that it was fixed
- Vendor documentation on model updates and their validation
The through-line: an AI coder is a tool operated under a human compliance program, not a replacement for one. The practice remains responsible for what it bills. Every framework decision should follow from that.
Ten questions to ask an AI coding vendor
- What exactly does your accuracy figure measure: primary code, full claim, modifiers, ICD-10 specificity?
- What was the ground truth standard, and who established it?
- Show me accuracy broken out by orthopaedic service line, including spine and revision arthroplasty.
- What is your undercoding rate versus your overcoding rate, reported separately?
- Is your confidence output calibrated, and can you show me the calibration curve?
- Who sets the human-review threshold, and can we change it?
- How do you handle the October 1 ICD-10 and January 1 CPT updates, and what is your validation process for each?
- What happens when a payer policy changes mid-year?
- Can we run an independent third-party audit of your output, and will you support it?
- What audit trail do you produce for a given claim? Can we see why a code was selected?
Question 10 is the one that separates products. An AI coder that cannot explain a code selection cannot be defended in an audit, regardless of how accurate it is.
Frequently asked questions
Is AI medical coding compliant with CMS requirements?
CMS does not prohibit the use of AI in coding and does not certify coding tools. The compliance obligation is unchanged: the practice must submit codes supported by the documentation in the medical record. What matters is that the practice maintains a documented governance program covering validation, human oversight, monitoring, and remediation, demonstrating how it ensures submitted codes are supported. AI is a tool operated under that program.
How accurate does an AI coder need to be?
The meaningful benchmark is not an absolute number but a comparison. Is it more accurate than your current process, measured by an independent auditor against the same documentation, broken out by service line and by error direction? Many orthopaedic practices discover their human baseline is lower than assumed, particularly on E/M leveling and modifier capture. Set your threshold against your measured baseline, not against a vendor's marketing figure.
Do we still need certified coders if we use AI coding?
Yes, and their role changes rather than shrinks. Certified coders move from producing codes to validating them, handling the low-confidence and complex cases the system routes to review, running the audit program, and owning payer-specific rules. This is generally higher-value work than volume coding, and it is the work that makes the AI defensible.
What is confidence calibration and why does it matter?
Calibration means the system's stated confidence matches its actual accuracy. When it says 95% confident, it is right roughly 95% of the time. It matters because the human review queue is triggered by low confidence. A miscalibrated, overconfident system pushes its errors past the safety net designed to catch them. Ask any vendor for a calibration curve, not just an accuracy number.
How often should we audit AI-generated codes?
More often than human coding, because AI behavior changes with model updates and code set changes while human error is comparatively stable. A practical cadence is weekly automated distributional monitoring, a monthly independent sample of 30 to 50 charts, a quarterly full audit by service line, and an event-triggered re-audit after any model update, annual code set change, or significant payer policy change.
What should an AI coding audit trail contain?
At minimum: the code selected, the confidence level, the specific documentation elements the system relied on, whether the case was routed to human review, the reviewer's determination and any override, and the version of the model that produced the output. If a vendor cannot produce a per-claim rationale, the output cannot be defended in a payer audit.
The practical case for auditability
There is a version of this conversation where auditing is framed as the price of adopting AI. That framing gets it backwards. Manual coding has never been meaningfully auditable at scale. You sample 30 charts a year out of tens of thousands and extrapolate. Autonomous coding produces a complete, structured, queryable record of every decision it made and why.
That is the first time an orthopaedic practice can actually answer the question "how are we coding?" rather than estimating it. Practices that treat the audit framework as a governance burden get a compliance file. Practices that treat it as a measurement system get a feedback loop, one that surfaces documentation gaps by surgeon, by service line, and by code, and closes them upstream.
Maia was built for that model. The AutoCoder operates inside the EHR, whether that is Athena, eClinicalWorks, Epic, ModMed, NextGen, or Tebra, recommending and populating codes, modifiers, and clinical justification before a human coder opens the chart, with the reasoning attached. Maia's auditing partnership with Karen Zupko & Associates exists specifically because independent validation is what makes autonomous coding defensible. Orthopaedic groups across the country have adopted it on that basis.
See how Maia's AutoCoder works automatically for orthopaedic practices. Book a demo at usemaia.com.




