
University of the Philippines (UP) Manila researchers revealed that ChatGPT falls short when compared to human neurosurgery residents, suggesting that the tool is not yet ready to support neurosurgical education despite having the capability of “understanding and generating natural language and other types of content.”
Results showed that ChatGPT’s accuracy ranged from 50.4 to 78.8 percent, while residents scored 58.3 to 73.7 percent, indicating that, overall, residents still outperformed the AI.
“Our meta-analysis showed that neurosurgery residents performed better than ChatGPT in answering neurosurgery board examination-like questions, although reviewed studies had high heterogeneity,” the study concludes.
In the study, the researchers reviewed published literature from November 2022, when ChatGPT was first released, up to October 2024, to evaluate its performance in specialized medical examinations.
The analysis covered six studies that compared ChatGPT’s performance with that of neurosurgery residents.
Although the overall results leaned in favor of the residents, the researchers noted that the studies they reviewed varied in their findings.
When the highest weighted study was removed, the results shifted to show better performance by ChatGPT.
“Further improvement is necessary before it can become a useful and reliable supplementary tool in the delivery of neurosurgical education,” the study added, highlighting the need for its continued development to be fully integrated into specialized medical training.
The study, titled “The performance of ChatGPT versus neurosurgery residents in neurosurgical board examination-like questions: a systematic review and meta-analysis,” sheds light on how artificial intelligence (AI) tools like OpenAI’s ChatGPT perform when presented with neurosurgery board examination-style questions.