Claimatix / Blog / Clinical AI
Physician judgment at machine speed: where AI belongs in workers' comp decisions
The law already answers the hardest question about AI in utilization review. The real work is designing around the answer.
By Ranjeet Randhawa, founder and CEO | September 22, 2026 | 10 min read

Public debate about artificial intelligence in utilization review is usually framed as a contest: will algorithms replace the physicians who decide whether treatment is medically necessary? In California workers' compensation, that question was settled long before large language models arrived.
The rule that already exists
Labor Code section 4610 provides that no one other than a licensed physician competent to evaluate the specific clinical issues involved may modify, delay or deny a request for authorization for reasons of medical necessity. It also requires that the criteria used be consistent with the Medical Treatment Utilization Schedule and disclosed when they form the basis of a modification or denial.[1]
Read as an engineering specification, that language is precise. It does not prohibit software. It defines where software must stop: at the decision to say no or not yet, and at the point where reasoning must be shown to the requesting physician and the injured worker.
Where regulation is heading
Two developments point the same direction, though neither was written for workers' compensation.
California's Physicians Make Decisions Act. Senate Bill 1120, signed in September 2024, governs how health plans and disability insurers use AI and algorithms in utilization review and requires that denials, delays or modifications based on medical necessity be made by a licensed physician or other qualified provider.[2] Legal analyses of the statute emphasize a second requirement: AI-supported determinations must be grounded in the individual patient's history and clinical circumstances, not group datasets alone.[3] The statute targets health coverage, but its logic mirrors section 4610.
The NAIC model bulletin. In December 2023 the National Association of Insurance Commissioners adopted a model bulletin on insurers' use of AI systems; by March 2025, 24 states had adopted it with little or no material change. It expects a written AI program covering governance, risk management, internal controls and vendor management, and readiness for regulatory inquiry into how AI is developed and used.[4] Several states apply it to every insurer they license.[5]
Common thread: AI may support the review, a qualified human decides, and the organization must be able to show how the whole thing works.
What the literature says about AI-assisted human review
The interesting risk is not that AI will make decisions by itself. It is that a physician who reviews twenty cases with a confident draft in front of them will defer to it.
Health informatics has a name for this: automation bias. A systematic review of decision support studies found that clinicians both accept incorrect system advice (errors of commission) and fail to act when a system stays silent (errors of omission), and that the mitigators are implementation factors — training, explicit accountability, how advice is displayed, and whether the system presents information rather than a recommendation.[6] A later review found that verification complexity matters: the harder it is for a user to independently check the advice, the more likely they are to simply follow it.[7]
The effect is measurable even among specialists. In a 2023 Radiology experiment, radiologists reading mammograms with AI-generated BI-RADS suggestions were significantly influenced by incorrect suggestions, with inexperienced readers most affected.[8]
Those findings have direct design implications for utilization review. A system that hands a reviewer a finished determination invites deference. A system that hands a reviewer the assembled evidence, the specific guideline section with a link to its source, and the gaps in the record, keeps verification cheap and judgment intact.
Where the judgment actually lives
California's IMR data shows where nuance concentrates. In 2025, pharmaceutical requests made up 31% of the treatment requests submitted to IMR, and opioids accounted for 22% of those. MTUS guidelines remained the primary resource for determining medical necessity, and the categories overturned most often were program services, behavioral and mental health services, and evaluations.[9]
Those high-overturn categories are precisely the ones where an individual worker's history, function and psychosocial context weigh most. That is a caution against automating the shortcut in exactly the places it is most tempting, and an argument for using automation to assemble the full picture faster so the reviewer's limited attention goes to the part only they can do.
It also matters that early decisions are consequential. Early opioid exposure after a back injury is associated with roughly double the risk of long-term work disability,[10] and non-indicated early MRI is associated with substantially longer disability and higher costs.[11] Speed without judgment is not neutral; it compounds.
A workable division of labor
Work the system should do. Log and timestamp every request on every channel. Check requests for completeness. Assemble the treatment history into a readable timeline. Surface the relevant guideline sections with citations the reviewer can open. Draft a rationale in the reviewer's own structure. Track every deadline and every notice.
Work that belongs to the physician. Weigh the evidence for this worker. Decide to approve, modify or deny. Conduct the peer-to-peer conversation. Own and sign the rationale.
Work that belongs to the record. Capture what the system prepared, what the physician changed, which guideline was relied on, and when each party was served. When an IMR reviewer, a judge or a market conduct examiner asks how a decision was made, the answer should be a record, not a reconstruction.
Five questions to ask any AI vendor
- Who makes the final call, and is that recorded? The system should make it structurally impossible to issue a modification or denial without a qualified physician's decision attached to it.
- Can every guideline citation be traced to its source? If a reviewer cannot open the exact MTUS section behind a draft in one click, verification is expensive and deference is the default.
- Is the analysis built from this worker's records? Population patterns can shape triage and priority. They cannot substitute for the individual file, and under SB 1120's logic they should not.
- What happens when evidence is missing? A well-built system flags the gap and generates a specific request for information. A poorly built one fills the gap with plausible language.
- What do you measure, and can I see it? Ask for the operational metrics that matter: how often the physician changes the draft, time-to-decision by review type, timeliness compliance, and the IMR overturn rate on decisions the system supported. A vendor who tracks only model accuracy is measuring the wrong system.
The takeaway
The law has answered who decides. The open question is whether the system around the reviewer produces a better decision faster, and whether it can prove it afterward. Designing for verification rather than convenience is the difference between AI that strengthens clinical review and AI that quietly replaces it while claiming not to.
This article is general information about workers' compensation rules and practice, not legal advice. Check current statutes, regulations and case law for your jurisdiction.
References
- California Labor Code § 4610. https://california.public.law/codes/labor_code_section_4610
- Office of Senator Josh Becker. "Governor signs Physicians Make Decisions Act, keeping medical decisions between patients and doctors, not AI." September 30, 2024. https://sd13.senate.ca.gov/news/press-release/september-30-2024/governor-signs-physicians-make-decisions-act-keeping-medical
- Kessenick Gamma LLP. "Understanding the Implications of California's SB 1120 for Healthcare Utilization Review." https://kessenick.com/implications-of-californias-sb-1120-for-healthcare-utilization/
- Quarles & Brady LLP. "Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers' Use of AI." 2025. https://quarles.com/newsroom/publications/nearly-half-of-states-have-now-adopted-naic-model-bulletin-on-insurers-use-of-ai
- National Association of Insurance Commissioners. "Implementation of NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers" (adoption map). https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121-127. https://doi.org/10.1136/amiajnl-2011-000089
- Lyell D, Coiera E. Automation bias and verification complexity: a systematic review. J Am Med Inform Assoc. 2017;24(2):423-431. https://doi.org/10.1093/jamia/ocw105
- Dratsch T, Chen X, Rezazade Mehrizi M, et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology. 2023;307(4):e222176. https://doi.org/10.1148/radiol.222176
- California Department of Industrial Relations. "DIR, DWC release Independent Medical Review report for 2025." 2026. https://dir.ca.gov/DIRNews/2026/2026-43.html
- Franklin GM, Stover BD, Turner JA, Fulton-Kehoe D, Wickizer TM. Early opioid prescription and subsequent disability among workers with back injuries. Spine. 2008;33(2):199-204. https://doi.org/10.1097/BRS.0b013e318160455c
- Webster BS, Bauer AZ, Choi Y, Cifuentes M, Pransky GS. Iatrogenic consequences of early magnetic resonance imaging in acute, work-related, disabling low back pain. Spine. 2013;38(22):1939-1946. https://pubmed.ncbi.nlm.nih.gov/23883826/