A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Huang, Yining; Tang, Keke; Chen, Meilian; Wang, Boyuan

Computer Science > Computation and Language

arXiv:2404.15777 (cs)

[Submitted on 24 Apr 2024 (v1), last revised 29 May 2024 (this version, v4)]

Title:A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Authors:Yining Huang, Keke Tang, Meilian Chen, Boyuan Wang

View PDF HTML (experimental)

Abstract:Since the inception of the Transformer architecture in 2017, Large Language Models (LLMs) such as GPT and BERT have evolved significantly, impacting various industries with their advanced capabilities in language understanding and generation. These models have shown potential to transform the medical field, highlighting the necessity for specialized evaluation frameworks to ensure their effective and ethical deployment. This comprehensive survey delineates the extensive application and requisite evaluation of LLMs within healthcare, emphasizing the critical need for empirical validation to fully exploit their capabilities in enhancing healthcare outcomes. Our survey is structured to provide an in-depth analysis of LLM applications across clinical settings, medical text data processing, research, education, and public health awareness. We begin by exploring the roles of LLMs in various medical applications, detailing their evaluation based on performance in tasks such as clinical diagnosis, medical text data processing, information retrieval, data analysis, and educational content generation. The subsequent sections offer a comprehensive discussion on the evaluation methods and metrics employed, including models, evaluators, and comparative experiments. We further examine the benchmarks and datasets utilized in these evaluations, providing a categorized description of benchmarks for tasks like question answering, summarization, information extraction, bioinformatics, information retrieval and general comprehensive benchmarks. This structure ensures a thorough understanding of how LLMs are assessed for their effectiveness, accuracy, usability, and ethical alignment in the medical domain. ...

Comments:	42 pages, 1 figure
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2404.15777 [cs.CL]
	(or arXiv:2404.15777v4 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2404.15777

Submission history

From: Yining Huang [view email]
[v1] Wed, 24 Apr 2024 09:55:24 UTC (36 KB)
[v2] Sun, 5 May 2024 16:44:58 UTC (55 KB)
[v3] Wed, 22 May 2024 08:57:19 UTC (77 KB)
[v4] Wed, 29 May 2024 15:50:43 UTC (477 KB)

Computer Science > Computation and Language

Title:A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators