Sep 30, 2026
ManyPress

Advertisement

Science

Researchers at the University of Wollongong found that generative AI models have significantly improved in legal reasoning, raising concerns about academic integrity in unsupervised assessments.

ManyPress

ManyPress

ManyPress Editorial

3 min readSource:Phys.org
Study Finds AI Models Outperforming Law Students in Recent University Exams

Key facts

  • •Researchers tested 18 AI-generated papers across criminal law and tort law subjects at the University of Wollongong.
  • •AI models outperformed 82.5% of students in criminal law and 61% of students in tort law.
  • •Seven of the 18 AI papers ranked at or above the 90th percentile of student performance.
  • •The study identified 'jagged capabilities,' noting that models could be excellent in one subject but weak in another.
  • •Researchers recommend a mix of supervised, AI-free exams and assessments that evaluate a student's ability to critically supervise AI output.

A new study from the University of Wollongong tested nine AI models across two compulsory law subjects, finding a significant improvement in performance compared to 2023. The AI tools, which had internet access but no subject-specific materials, outperformed the majority of students in both criminal law and tort law exams. Researchers noted that while AI capabilities have advanced, the models still exhibit inconsistent performance and occasional errors, prompting calls for new university assessment strategies.

By the numbers

Average score for AI in criminal law76.3%
Average score for AI in tort law66%
Percentage of students outperformed by AI in criminal law82.5%
Percentage of students outperformed by AI in tort law61%

Significant Gains in Legal Reasoning

In the 2023 criminal law exam, AI models averaged 52.5% and performed at the 22nd percentile. In the new study, AI papers in criminal law averaged 76.3%, outperforming 82.5% of students. For tort law, the models averaged 66%, outperforming 61% of students. Seven of the 18 AI-generated papers ranked at or above the 90th percentile of student results.

Inconsistencies and Hallucinations

Despite the improved grades, the study found that AI models remain unreliable legal experts. Performance varied between subjects, with some models showing strong analysis alongside poor citations or fabricated authorities. Researchers observed that while some models had lower rates of hallucination, they also cited fewer sources overall.

Proposed Assessment Changes

The researchers suggest that universities move away from reliance on unsupervised assessments. They propose a three-pronged approach: maintaining AI-free, monitored in-person exams; requiring students to collaborate with AI while grading the process rather than just the final product; and using a 'relay' model where students alternate between AI-assisted drafting and independent, supervised critique of that work.

Timeline

  1. End of 2022
    ChatGPT was released, prompting university concerns regarding AI-assisted cheating.
  2. 2023
    Researchers conducted an initial experiment testing AI on a criminal law exam.
  3. 2026
    The study 'Legal Minds Vs. Neural Networks' was published in Law, Technology and Humans.

Advertisement

This article was independently rewritten by ManyPress editorial AI from reporting originally published by Phys.org.

Science