沈阳医学院 · 辽宁 · PDF · 11 页 · 10185KB
View Online Export Citation RESEARCH ARTICLE | MARCH 04 2026 An efficient YOLOv12n model is used for multi-scale brain tumor detection Lan Liu ; Ronghuan Wu ; Chengyang Zhang; Jianhui Xu AIP Advances 16, 035309 (2026) https://doi.org/10.1063/5.0321006 Articles You May Be Interested In Efficient rice pest intelligent detection framework for field environments AIP Advances (October 2025) Construction of an efficient blood cell detection model for small targets and overlapping areas AIP Advances (September 2025) A photovoltaic panel defect detection framework enhanced by deep learning AIP Advances (September 2025) 0 6 M a y 2 0 2 6 0 2 : 0 4 : 1 6 AIP Advances ARTICLE pubs.aip.org/aip/adv An efficient YOLOv12n model is used for multi-scale brain tumor detection Cite as: AIP Advances 16, 035309 (2026); doi: 10.1063/5.0321006 Submitted: 5 January 2026 • Accepted: 8 February 2026 • Published Online: 4 March 2026 Lan Liu,1 Ronghuan Wu,2 Chengyang Zhang,3 and Jianhui Xu4,a) AFFILIATIONS 1 Department of Basic Medicine, Changde Vocational Technical College, Hunan, Changde, 415000, China 2China Mobile Guizhou Co., Ltd. Guiyang Branch, Guiyang 550001, China 3Amazon, 410 Terry Avenue North, Seattle, Washington 98109, USA 4School of Medical Information Engineering, Shenyang Medical College, Shenyang, Liaoning 110043, China a)Author to whom correspondence should be addressed: 329954097@qq.com ABSTRACT Accurate detection of brain tumors is critical for early diagnosis and treatment planning, yet it remains challenging due to the irregular shapes, blurred boundaries, and low contrast of lesions, particularly in Gliomas. To address the limitations of existing methods in balancing accuracy with computational efficiency, this paper proposes a lightweight brain tumor detection framework based on an improved YOLOv12n. We introduce three core innovations to enhance feature representation: (1) the A2C2f-CGLU-DYT module, which integrates dynamic activation functions to strengthen local nonlinear modeling and adaptability to input variations; (2) the C2TSSA module, which utilizes token statistics self-attention to efficiently capture global long-range dependencies without the heavy computational cost of traditional transformers; and (3) the CGAFusion module, which employs content-guided attention to effectively fuse shallow geometric details with deep semantic features. Experimental results on the Brain Tumor dataset demonstrate that the proposed method achieves a precision of 92.7%, a recall of 87.4%, and an mAP@0.5 of 93.8%, outperforming the baseline YOLOv12n by 0.7%, 0.9%, and 1.2%, respectively. Notably, in the challenging Glioma category, the model attains an accuracy of 84.6%, significantly surpassing YOLOv10n by a margin of 8.6%. Furthermore, validation on the Blood Cell Count dataset yields an mAP@0.5 of 94.1%, confirming the model’s robust cross-dataset generalization ability. © 2026 Author(s). All article content, except where otherwise noted, is licensed under a Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/). https://doi.org/10.1063/5.0321006 I. INTRODUCTION Brain tumors are a serious threat to human life and health, being a type of central nervous system disease with high disability and mortality rates.1 Clinical studies have shown that early diagnosis and timely treatment are crucial for improving patient survival rates and prognosis.2 Medical imaging technology, particularly Magnetic Resonance Imaging (MRI), with its advantages of non-invasiveness, high resolution, and multi-parameter imaging,3 has become a core tool for clinical diagnosis, classification, and treatment planning for brain tumors. However, traditional manual image reading methods not only rely on the professional experience of doctors but also suffer from problems such as high subjectivity, poor reproducibility, and heavy workload, facing challenges in efficiency and accuracy when processing large volumes of m