Skip to main content

BUS-Large: A large harmonized and curated breast ultrasound dataset with segmentation masks and clinical labels

https://doi.org/10.71775/kth.frfd6-ymw69
Breast Ultrasound (BUS) imaging plays an important role in the early detection and diagnosis of breast cancer. The rapid advancement of deep learning-based Computer-Aided Diagnosis (CAD) systems has demonstrated remarkable potential. However, the clinical applicability of CAD is constrained by the scarcity of large-scale, diverse, and well-annotated datasets. Existing public BUS datasets are highly fragmented. To address these gaps, we present BUS-Large, a large-scale curated BUS dataset for method development and evaluation, comprising more than 21,00 images from 21 publicly available datasets across multiple countries and clinical institutions. The dataset integrates all the public available static image datasets. All cohorts have been harmonized with unified diagnostic labels, pixel-level segmentation mask, and standardized BI-RADS clinical information. To the best of our knowledge, this is one of the most extensive datasets providing lesion segmentation masks annotated from publicly available BUS datasets.  This work was supported by grants from Marie Skłodowska-Curie Doctoral Networks Actions (HORIZON-MSCA-2021-DN-01-01;  101073222), Cancerfonden (22-2389 Pj). In addition to the MAIA platform at KTH, the computations were enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Academic Infrastructure for Supercomputing in Sweden.  We also acknowledge the public dataset owners for making the imaging and clinical data used in this study publicly available.
Go to data source
https://doi.org/10.71775/kth.frfd6-ymw69

Citation and access

Administrative information

Identifiers

Funding

Topic and keywords

Relations

Metadata

kth-datarepository
kth