ChatGPT Enters Courtroom Expert Evidence: Analysing AI’s Reliability and Procedural Fairness in Australian Law – Gauci v Roo [2024] NSWDC 361.

“…[An expert’s] resort to AI, Chat-GPT suggests to me that, in the absence of evidence, I ought not infer expertise in radiology in those without qualifications in that discipline.”

Ad/Marketing Communication

This legal article/report forms part of my ongoing legal commentary on the use of artificial intelligence within the justice system. It supports my work in teaching, lecturing, and writing about AI and the law and is published to promote my practice. Not legal advice. Not Direct/Public Access. All instructions via clerks at Doughty Street Chambers. This legal article concerns AI Law.

In Gauci v Roo [2024] NSWDC 361, the District Court of New South Wales tackled a fascinating interplay between medical evidence and procedural fairness.

At its core, the case raised questions about the reasonableness of subjecting a young plaintiff to further radiological testing, but it also revealed how AI, specifically ChatGPT, had found its way into the medical opinion process, something which, anecdotally, I have heard is happening more frequently. This case however was different from a case I have discussed previously where it was suggested Chat GPT might be analogous to expert evidence.

So, I will discuss what this case was briefly about before considering AI’s role and future implications.

Background and Key Details

  • The plaintiff was injured in a 2018 road accident while riding a push bike. He brought a compensation claim against the defendant, seeking damages under the Motor Accident Injuries Act 2017 (NSW).
  • An important issue was whether the defendant could compel the plaintiff to undergo a further CT scan of his femur in order to determine malrotation. This issue integral to establishing the plaintiff’s permanent impairment rating.
  • The plaintiff had already undergone various scans, including a CT scan and X-rays, which had produced inconsistent findings about malrotation. Some doctors claimed imaging was the best evidence; others relied on clinical examinations.
  • Ultimately, the Court dismissed the defendant’s application to force another CT scan. Balancing the potential risks of radiation exposure and the availability of alternative methods to assess malrotation, the judge concluded it was not “reasonable” to require further scanning.

How AI Featured in the Case

Although the Court’s ruling itself did not hinge on advanced machine learning or automated decision-making, AI did make a cameo appearance in the form of an orthopaedic surgeon consulting ChatGPT. In one of his reports, the surgeon explicitly mentioned using ChatGPT to look up radiation levels associated with a specific protocol (the “Perth protocol”) for CT assessment. This doctor then cited the AI-generated information in forming his view that the radiation exposure from the proposed scan was not trivial.

Catsanos SC DCJ observed at paragraph 52:

“52 Other than Dr Steinberg, I have no evidence from a radiologist. I am left with the evidence of Doctors Machart and Gehr, being orthopaedic surgeons, and Dr Gorman, a medical assessor, but not a radiologist. Those doctors are no doubt familiar with the use of radiological investigations for diagnostic purposes and perhaps have generalised medical knowledge about risks associated with those investigations. However, there is no evidence before me that they have any particular expertise in the administration of radiological investigations, nor safety considerations associated with levels and frequency of exposure to radiation. Indeed, Dr Gehr’s resort to AI, Chat-GPT  suggests to me that, in the absence of evidence, I ought not infer expertise in radiology in those without qualifications in that discipline.”

Implications and Reflection

Gauci v Roo demonstrates that even routine requests, like asking someone to have another CT scan, can highlight bigger, more modern issues. The indirect involvement of AI here makes us think about how our legal system is dealing with new technology, how trustworthy AI-generated information really is, and whether it’s right to rely on AI when giving expert opinions.

This case shows that expert evidence and the ways professionals gather information are changing. I was happy to see the expert openly admitting he used ChatGPT. In other cases, it’s been disheartening to watch judges struggle with uncertainty about whether the information presented to them was AI-generated. This openness also raises important concerns. When an expert relies on AI, like ChatGPT, to determine crucial details, such as radiation dosage, how accurate or reliable is that information? Courts may become more cautious about experts who depend on AI instead of their own experience or established research.

There are also important issues around fairness and transparency. Should experts be required to clearly state exactly what AI tools they’ve used in their research and when drafting their final reports? Could someone argue it’s unfair if the opposing side uses AI-generated information that’s difficult to verify?

What about the sophistication of AI tools themselves? Are experts relying on basic versions of ChatGPT inherently less reliable than those with access to more advanced, professional AI services?

With the growing use of AI, courts might need clear guidelines on when and how to trust AI-generated information. I can imagine arguments suggesting AI might lower the standards expected from expert witnesses, but also arguments supporting the exact opposite. Surely, the ideal scenario would be a symbiotic relationship between expert witnesses and sophisticated LLMs?

Could AI tools like ChatGPT ever become as respected, or even more respected, than human experts on certain topics? My tentative answer is yes, but only if developers manage to overcome current issues like bias and deceptive outputs in the AI reasoning process. Otherwise, LLMs will likely remain subject to criticism and judicial scepticism.