Loading…

Readability and Appropriateness of Responses Generated by ChatGPT 3.5, ChatGPT 4.0, Gemini, and Microsoft Copilot for FAQs in Refractive Surgery

To assess the appropriateness and readability of large language model (LLM) chatbots' answers to frequently asked questions about refractive surgery. Four commonly used LLM chatbots were asked 40 questions frequently asked by patients about refractive surgery. The appropriateness of the answers...

Full description

Saved in:
Bibliographic Details
Published in:Turkish journal of ophthalmology 2024-12, Vol.54 (6), p.313
Main Authors: Aydın, Fahri Onur, Aksoy, Burakhan Kürşat, Ceylan, Ali, Akbaş, Yusuf Berk, Ermiş, Serhat, Kepez Yıldız, Burçin, Yıldırım, Yusuf
Format: Article
Language:English
Subjects:
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:To assess the appropriateness and readability of large language model (LLM) chatbots' answers to frequently asked questions about refractive surgery. Four commonly used LLM chatbots were asked 40 questions frequently asked by patients about refractive surgery. The appropriateness of the answers was evaluated by 2 experienced refractive surgeons. Readability was evaluated with 5 different indexes. Based on the responses generated by the LLM chatbots, 45% (n=18) of the answers given by ChatGPT 3.5 were correct, while this rate was 52.5% (n=21) for ChatGPT 4.0, 87.5% (n=35) for Gemini, and 60% (n=24) for Copilot. In terms of readability, it was observed that all LLM chatbots were very difficult to read and required a university degree. These LLM chatbots, which are finding a place in our daily lives, can occasionally provide inappropriate answers. Although all were difficult to read, Gemini was the most successful LLM chatbot in terms of generating appropriate answers and was relatively better in terms of readability.
ISSN:2149-8709
2149-8709
DOI:10.4274/tjo.galenos.2024.28234