<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <atom:link rel="self" type="application/rss+xml" href="https://researchdata.se/sv/catalogue/search.rss?freeKeyword=Ecology+not+elsewhere+classified"/>
    <link>https://researchdata.se/sv/catalogue</link>
    <title>Researchdata.se</title>
    <description>Search results</description>
    <language>sv</language>
    <item>
      <title>Baltic coastal meadow specialist plant abundances and habitat extent variables in the 1960s and 2024</title>
      <description>This data consists of an inventory and a re-inventory of plant specialist abundances in Baltic coastal meadows, together with explanatory variables used to model species occurrence and abundance changes between the time steps. The original data was collected by Germund Tyler and published in Tyler (1969) as maps. This original data was collected from 76 Baltic coastal meadows, but here we only include the 65 for which we have data from the re-inventory of 2024. Four of the eleven meadows excluded from this dataset were situated on inaccessible islands, two have been converted to parking lots and five could not be located for the re-inventory. Habitat size, habitat amount, number of coastal meadows and management were calculated using a time series of aerial images. The management status was determined for all sites in each time step in the aerial image time series; all managed coastal meadows were digitized for the aerial images from the 1960s and 2023. See the belonging paper for more information on the plant inventories and the aerial image interpretation.

Reference to the original dataset: Tyler, G. (1969). Studies in the ecology of Baltic sea-shore meadows. 2. Flora and vegetation. Opera Botanica a Societate Botanica Lundensi, 25.</description>
      <pubDate>Tue, 27 May 2025 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-29117846</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-29117846</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Lukas Rimondini</dc:creator>
    </item>
    <item>
      <title>Pre-release of v1 for the Processed ASV data from the Insect Biome Atlas Project.</title>
      <description>The data was collected and processed in the context of the Swedish Insect Biome Atlas (https://www.insectbiomeatlas.org/)  (IBA). This dataset is a pre-release of v1 for the Processed ASV data from the Insect Biome Atlas Project. 

For more detailed information see Processed ASV data from the Insect Biome Atlas Project: https://doi.org/10.17044/scilifelab.27202368.v1

References:

- Iwaszkiewicz-Eggebrecht, E., Łukasik, P., Buczek, M., Deng, J., Hartop, E. A., Havnås, H., ... &amp; Miraldo, A. (2023). FAVIS: Fast and versatile protocol for non-destructive metabarcoding of bulk insect samples. PloS one, 18(7), e0286272.
- Miraldo, A., Iwaszkiewicz-Eggebrecht, E., Sundh, J., Manoharan, L., Granqvist, E., Andersson, A., Łukasik, P., Roslin, T., Tack, A. J. M., &amp; Ronquist, F. (2024). Amplicon sequence variants from the Insect Biome Atlas project (Version 1). SciLifeLab. https://doi.org/10.17044/scilifelab.25480681.v1
- Sundh, J. (2022). COI reference sequences from BOLD DB (Version 4). SciLifeLab. https://doi.org/10.17044/scilifelab.20514192.v4</description>
      <pubDate>Thu, 23 Jan 2025 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-28217918</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-28217918</guid>
      <dc:publisher>Uppsala universitet</dc:publisher>
      <dc:creator>Adrian Baggström</dc:creator>
      <dc:creator>Robert Goodsell</dc:creator>
      <dc:creator>Laura van Dijk</dc:creator>
      <dc:creator>Elzbieta Iwaszkiewicz-Eggebrecht</dc:creator>
      <dc:creator>Andreia Miraldo</dc:creator>
      <dc:creator>Ayco J. M. Tack</dc:creator>
      <dc:creator>Tobias Andermann</dc:creator>
    </item>
    <item>
      <title>Amplicon sequence variants from the Insect Biome Atlas project</title>
      <description>General informationThe Insect Biome Atlas project was supported by the Knut and Alice Wallenberg Foundation (dnr 2017.0088). The project analyzed the insect faunas of Sweden and Madagascar, and their associated microbiomes, mainly using DNA metabarcoding of Malaise trap samples collected in 2019 (Sweden) or 2019–2020 (Madagascar).

Please cite this version of the dataset as: Miraldo A, Iwaszkiewicz-Eggebrecht E, Sundh J, Lokeshwaran M, Granqvist E, Andersson AF, Lukasik P, Roslin T, Tack A, Ronquist F. 2024. Dataset of amplicon sequence variants (ASVs) from the Insect Biome Atlas Project, version 5. https://doi.org/10.17044/scilifelab.25480681

Dataset descriptionThis dataset contains amplicon sequence variants (ASVs) generated from high-throughput sequencing of the cytochrome c oxidase subunit I (COI) gene from Malaise trap samples (lysates, homogenates and preservative ethanol) and soil and litter samples. It includes ASV sequences and abundance information (number of reads) as well as metadata files that are needed to interpret and analyse the data further. Future versions of the dataset will include additional data. NB! All ASV files include ASVs that represent biological and synthetic spike-ins.

MethodsSamples were sequenced using Illumina technology. Raw data are available at the European Nucleotide Archive (ENA) under project PRJEB61109. The raw sequence data was preprocessed using a Snakemake workflow (https://github.com/biodiversitydata-se/amplicon-multi-cutadapt) . Preprocessed reads were then used as input to the AmpliSeq (https://github.com/nf-core/ampliseq)  Nextflow (v.2.1.0) pipeline to generate ASVs.

Available dataTwo types of files are provided: ASV files and metadata files. Files marked with 'SE' and 'MG' contain data from Sweden and Madagascar, respectively.

The file shasum.txt contains checksums for each of the files.

ASV filesASV sequences in fasta format are found in files CO1_asv_seqs_SE.fasta.gz and CO1_asv_seqs_MG.fasta.gz. Counts of ASVs in each sample are in CO1_asv_counts_SE.tsv.gz and CO1_asv_counts.MG.tsv.gz. The Swedish dataset contains 821,559 ASVs in 6,169 samples. The Madagascar dataset contains 701,769 ASVs in 2,286 samples.

Metadata filesFour types of metadata files are included:

- sequencing_metadata files with information about samples that were processed in the lab and sequenced
- samples_metadata files with information about samples that were collected in the field.
- sites_metadata files with information about sites where samples were collected.
- sipke-ins metadata files with information about spike-ins added to each malaise trap sample at the time of sample processing in the lab.
Sequencing metadata filesThe two sequencing metadata files CO1_sequencing_metadata_SE.tsv and CO1_sequencing_metadata_MG.tsv contain information about samples that were sequenced. For details on the columns of these files, see the README.txt file.

Samples metadata filesFour samples_metadata files are included in this dataset with information about each sample that was collected in the field. For samples collected with malaise traps we have two files, one for each country: samples_metadata_malaise_SE.tsv and samples_metadata_malaise_MG.tsv. See the README.txt file for details about the columns of these files.

For arthropod samples collected from litter and soil we have two files, one for each country: samples_metadata_soil_litter_SE.tsv and samples_metadata_litter_MG.tsv. Note that for Madagascar we did not collect arthropod samples from soil. Also note that for Madagascar we collected four leaf litter samples at each trap location, one sample in each direction of the Malaise trap (front, back, left and right); whilst for Sweden we collected only one sample at each trap location. For details on the columns of these files, see the README.txt file.

Sites metadata filesThere are two files that contain information about sampling sites, one for each country: sites_metadata_SE.tsv and sites_metadata_MG.tsv. See the README.txt file for more information.

Spike-ins metadata filesWe provide three files with information about spike-ins used when processing samples in the lab: biological_spikes_taxonomy_SE.tsv and biological_spikes_taxonomy_MG.tsv contain taxonomic information on biological spike ins while the file synthetic_spikes_info.tsv has information on synthetic spike ins. See README.txt for more information.

Other complementary data filesWe present complementary data on soil chemistry collected at each sampling location in both Sweden and Madagascar, stand characteristics collected at each sampling location in Madagascar and biomass/count data for a selected number of malaise trap samples from the Insect Biome Atlas project (n=24) and the Swedish Insect Inventory Project (n=224).

Soil chemistry dataWe provide two datasets, one for each country, on soil chemistry (soil_chemistry_SE.tsv and soil_chemistry_MG.tsv) that store information on soil nutrients from soil samples collected at the same sampling sites as the arthropod communities. Topsoil (0-20cm) was sampled at 5 sites around each Malaise trap in both Sweden and Madagascar: one soil core (6 cm diameter) at the center of trap and one soil core on each of the four “sides” of the trap five meters away from the trap. Soil samples at each site were taken as composite samples from the five locations. Soil samples collected in Sweden were analysed at Eurofins in Sweden and the ones collected in Madagascar were analysed at the Laboratoire des Radioisotopes in Madagascar. As samples from each country were analysed at different laboratories the variables on soil nutrients presented in each dataset differ slightly. See the README.txt file for more information on the columns of each of these files.

Stand characteristics dataStanding characteristics were only measured in Madagascar as extensive data on landscape composition and vegetation structure at the sampling sites in Sweden had already been compiled as part of the National Inventory of Landscapes in Sweden (NILS) and data are publicly available here (https://www.slu.se/en/Collaborative-Centres-and-Projects/nils/nils-inventory-2003-2020/) .

The file stand_characteristics_MG.tsv contains information on a set of standing characteristics from Madagascar related to tree density (DBH, shading, etc). Information about columns in this file is found in the README.txt file.

Biomass and count dataTo allow an assessment of how the biomass of a Malaise trap sample translates to the number of specimens, we provide two files describing samples from Sweden, for which we measured the biomass and also counted all the specimens in the sample. The first set comprises 24 samples from the IBA field campaign (biomass_count_IBA.tsv), and the second set comprises 224 samples from a separate Swedish Malaise trapping campaign (Swedish Insect Inventory Project) in 2018–2019 (biomass_count_SIIP.tsv). For the latter dataset, we provide the site and sample metadata in the same file. Details about columns in these files are found in the README.txt file.

References:

Egnér, H., Riehm, H., &amp; Domingo, W. (1960). Untersuchungen über die chemische Bodenanalyse als Grundlage für die Beurteilung des Nährstoffzustandes der Böden. II. Chemische Extraktionsmethoden zur Phosphor-und Kaliumbestimmung. Kungliga Lantbrukshögskolans Annaler, 26, 199–215.</description>
      <pubDate>Mon, 25 Nov 2024 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-25480681</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-25480681</guid>
      <dc:publisher>Naturhistoriska riksmuseet</dc:publisher>
      <dc:creator>Andreia Miraldo</dc:creator>
      <dc:creator>Elzbieta Iwaszkiewicz-Eggebrecht</dc:creator>
      <dc:creator>John Sundh</dc:creator>
      <dc:creator>Lokeshwaran Manoharan</dc:creator>
      <dc:creator>Emma Granqvist</dc:creator>
      <dc:creator>Anders Andersson</dc:creator>
      <dc:creator>Piotr Łukasik</dc:creator>
      <dc:creator>Tomas Roslin</dc:creator>
      <dc:creator>Ayco J. M. Tack</dc:creator>
      <dc:creator>Fredrik Ronquist</dc:creator>
    </item>
    <item>
      <title>Processed ASV data from the Insect Biome Atlas Project</title>
      <description>NOTE: The gzipped files in this upload have mistakenly been compressed twice. To decompress the files, please run e.g.:
gunzip -c CO1_asv_counts_SE.tsv.gz | gunzip -c &gt; CO1_asv_counts_SE.tsv



The Insect Biome Atlas project was supported by the Knut and Alice Wallenberg Foundation (dnr 2017.0088). The project analyzed the insect faunas of Sweden and Madagascar, and their associated microbiomes, mainly using DNA metabarcoding of Malaise trap samples collected in 2019 (Sweden) or 2019–2020 (Madagascar).

Please cite this version of the dataset as: Miraldo A, Iwaszkiewicz-Eggebrecht E, Sundh J, Lokeshwaran M, Granqvist E, Goodsell R, Andersson AF, Lukasik P, Roslin T, Tack A, Ronquist F. 2024. Processed ASV data from the Insect Biome Atlas Project, version 3. doi:10.17044/scilifelab.27202368.v3 or https://doi.org/10.17044/scilifelab.27202368.v3

This dataset contains the results from bioinformatic processing of version 1 of the amplicon sequence variant (ASV) data from the Insect Biome Atlas project (Miraldo et al. 2024), that is, the cytochrome oxidase subunit 1 (CO1) metabarcoding data from Malaise trap samples processed using the FAVIS mild lysis protocol (Iwaszkiewicz et al. 2023). The bioinformatic processing involved: (1) taxonomic assignment of ASVs, (2) chimera removal; (3) clustering into OTUs; (4) noise filtering and (5) cleaning. The clustering step involved resolution of the taxonomic annotation of the cluster and identification of a representative ASV. The noise filtering step involved removal of ASV clusters identified as potentially originating from nuclear mitochondrial DNA (NUMTs) or representing other types of error or noise. The cleaning step involved removal of ASV clusters present in &gt;5% of negative control samples. ASV taxonomic assignments, ASV cluster designations, consensus taxonomies and summed counts of clusters in the sequenced samples are provided in compressed tab-separated files. Sequences of cluster representatives are provided in compressed FASTA format files. The bioinformatic processing pipeline is further described in Sundh et al. (2024). NB! All result files include ASVs and clusters that represent biological and synthetic spike-ins.

MethodsTaxonomic assignmentASVs were taxonomically assigned using kmer-based methods implemented in a Snakemake workflow available here (https://github.com/insect-biome-atlas/happ) . Specifically ASVs were assigned a taxonomy using the SINTAX algorithm in vsearch (v2.21.2) using a CO1 database constructed from the Barcode Of Life Data System (Sundh 2022). ASVs assigned to Class 'Insecta' or 'Collembola' but unassigned at lower taxonomic ranks were then placed into a reference phylogeny of 49,325 insect species (represented by 49,338 sequences) using the phylogenetic placement tool EPA-NG with subsequent taxonomic assignments using GAPPA. Assignments at the order level in this second pass were used to update the first kmer-based assignments, but only at the order level, leaving child ranks with the ‘unclassified’ prefix.

Chimera removalThe workflow first identifies chimeric ASVs in the input data using the ‘uchime_denovo’ method implemented in vsearch. This was done with a so-called ‘strict samplewise’ strategy where each sample was analysed separately (hence the ‘samplewise’ notation), only comparing ASVs present in the same sample. Further, ASVs had to be identified as chimeric in all samples where they were present (corresponding to the ‘strict’ notation) in order to be removed as chimeric.

ASV clusteringNon-chimeric sequences were then split by family-level taxonomic assignments and ASVs within each family were clustered in parallel using swarm (v3.1.0) with differences=15. Representative ASVs were selected for each generated cluster by taking the ASV with the highest relative abundance across all samples in a cluster. Counts were generated at the cluster level by summing over all ASVs in each cluster.

Consensus taxonomyA consensus taxonomy was created for each cluster by taking into account the taxonomic assignments of all ASVs in a cluster as well as the total abundance of ASVs. For each cluster, starting at the most resolved taxonomic level, each unique taxonomic assignment was weighted by the sum of read counts of ASVs with that assignment. If a single weighted assignment made up 80% or more of all weighted assignments at that rank, that taxonomy was propagated to the ASV cluster, including parent rank assignments. If no taxonomic assignment was above the 80% threshold, the algorithm continued to the parent rank in the taxonomy. Taxonomic assignments at any available child ranks were set to the consensus assignment prefixed with ‘unresolved’.

Noise filtering and cleaningThe clustered data was further cleaned from NUMTs and other types of noise using the NEEAT algorithm, which takes taxonomic annotation, correlations in occurrence across samples (‘echo signal’) and evolutionary signatures into account, as well as cluster abundance (Sundh et al., 2024). We used default settings for all parameters in the evolutionary and distributional filtering steps, and removed clusters unassigned at the order level and with less than 3 reads summed across each dataset.

As a last clean-up step in the noise filtering, clusters containing at least one ASV present in more than 5% of blanks were removed. Further, we removed ASvs assigned to a reference sequence in the BOLD database annotated as Zoarces gillii (BOLD:AEB5125), a fish found between Japan and eastern Korea. Closer inspection revealed that this was a mis-annotated bacterial sequence and ASVs assigned to this reference most likely represent bacterial sequences in our dataset. This record has been deleted from BOLD after our custom reference database was constructed.

The chimera filtering and ASV clustering methods have been implemented in a Snakemake workflow available here (https://github.com/insect-biome-atlas/happ) . This workflow takes as input:

- The ASV sequences in FASTA format
- A tab-delimited file of counts of ASVs (rows) in samples (columns)
Data for 1) and 2) are available at https://doi.org/10.17044/scilifelab.25480681.v5

Cleaning of ASV clusters in controls and identification of spikeins was done with a custom R script available here (https://github.com/insect-biome-atlas/utils) .

Available dataProcessed ASV data filesASV taxonomic assignments, non-chimeric ASV cluster designations, consensus taxonomies, sequences of cluster representatives and summed counts of clusters in the sequenced samples are provided in compressed tab-separated files. Files are organized by country (Sweden and Madagascar), marked by the suffixes SE and MG, respectively.

Taxonomic assignmentsThe files asv_taxonomy_[SE|MG].tsv.gz are tab-separated files with taxonomic assignments using SINTAX+EPA-NG for all ASVs. Columns:

- ASV: The id of the ASV
- Kingdom, Phylum, Class, Order, Family, Genus, Species, BOLD_bin: Taxonomic assignment for each rank.
If an ASV was unclassified at a particular rank, the taxonomic label is prefixed with ‘unclassified.’ followed by the taxonomic assignment of the most resolved parent rank.

The files asv_taxonomy_sintax_[SE|MG].tsv.gz, asv_taxonomy_epang_[SE|MG].tsv.gz and asv_taxonomy_vsearch_[SE|MG].tsv.gz have the same structure, but contain results from assignments with SINTAX, EPA-NG and VSEARCH, respectively.

Cluster assignmentsThe files cluster_taxonomy_[SE|MG].tsv are tab-separated files containing all non-chimeric ASVs (that is, the ASVs passing the chimera-filtering step) with their corresponding taxonomic and cluster assignments. Columns:

- ASV: ASV id
- cluster: name of designated cluster
- median: the median of normalized reads across all samples for each ASV
- Kingdom, Phylum, Class, Order, Family, Genus, Species, BOLD_bin: taxonomic assignment of each ASV
- representative: contains 1 if ASV is a representative of its cluster, otherwise 0
Cluster countsThe files cluster_counts_[SE|MG].tsv are tab-separated files with read counts of ASV clusters (rows) in samples (columns). Counts have been summed for all ASVs belonging to each cluster. Note that these files contain counts for biological spike-ins and for Sweden also synthetic spike-ins.

Sequences of cluster representativesThe files cluster_reps_[SE|MG].fasta are text files in FASTA format with representative sequences for each cluster. The fasta headers have the format “&gt;ASV_ID CLUSTER_NAME”.

Consensus taxonomyThe files cluster_consensus_taxonomy_[SE|MG].tsv are tab-separated files with consensus taxonomy of each generated ASV cluster. Columns are the same as in asv_taxonomy_[SE|MG].tsv.

Noise-filtered dataThe files prefixed with 'noise_filtered' contain data that has been cleaned from NUMTs and other types of noise using the NEEAT algorithm. The files contain the same information as the cluster files, but only for clusters that passed the noise filtering step.

Cleaned noise filtered dataThe files prefixed with 'cleaned_noise_filtered' contain data that has been cleaned from NUMTs and other types of noise using the NEEAT algorithm, and further cleaned from clusters present in &gt;5% of blanks. The files contain the same information as the cluster files, but only for clusters that passed the noise filtering and cleaning steps.

Additional filesThe files removed_control_tax_[SE|MG].tsv.gz contain the ASV clusters removed from each dataset as part of cleaning.

The files spikeins_tax_[SE|MG].tsv.gz contain the taxonomic assignments of the biological spike-ins identified.

References:

- Iwaszkiewicz-Eggebrecht, E., Łukasik, P., Buczek, M., Deng, J., Hartop, E. A., Havnås, H., ... &amp; Miraldo, A. (2023). FAVIS: Fast and versatile protocol for non-destructive metabarcoding of bulk insect samples. PloS one, 18(7), e0286272.
- Miraldo, A., Iwaszkiewicz-Eggebrecht, E., Sundh, J., Manoharan, L., Granqvist, E., Andersson, A., Łukasik, P., Roslin, T., Tack, A. J. M., &amp; Ronquist, F. (2024). Amplicon sequence variants from the Insect Biome Atlas project (Version 5). SciLifeLab. https://doi.org/10.17044/scilifelab.25480681.v5
- Sundh, J. (2022). COI reference sequences from BOLD DB (Version 4). SciLifeLab. https://doi.org/10.17044/scilifelab.20514192.v4</description>
      <pubDate>Wed, 30 Oct 2024 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-27202368</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-27202368</guid>
      <dc:publisher>Naturhistoriska riksmuseet</dc:publisher>
      <dc:creator>Andreia Miraldo</dc:creator>
      <dc:creator>Elzbieta Iwaszkiewicz-Eggebrecht</dc:creator>
      <dc:creator>John Sundh</dc:creator>
      <dc:creator>Lokeshwaran Manoharan</dc:creator>
      <dc:creator>Emma Granqvist</dc:creator>
      <dc:creator>Anders F. Andersson</dc:creator>
      <dc:creator>Piotr Łukasik</dc:creator>
      <dc:creator>Tomas Roslin</dc:creator>
      <dc:creator>Ayco J. M. Tack</dc:creator>
      <dc:creator>Fredrik Ronquist</dc:creator>
    </item>
    <item>
      <title>Data from: Reproductive success, fruit removal and local distribution patterns in the early-flowering shrub Daphne mezereum</title>
      <description>This is the data from the article 

Reproductive success, fruit removal and local distribution patterns in the early-flowering shrub Daphne mezereum

(DOI: XXXXXX)

Matilda Arnell, Ove Eriksson and Johan Ehrlén

DESCRIPTION

In this study we mapped the spatial distribution of individuals in a population of the early flowering, fleshy-fruited shrub Daphne mezereum, in a forest in boreo-nemoral Sweden for three consecutive years (2016-2018). For all mapped individuals we collected data on numbers of flowers and fruits and fruit removal. In one year, 2019, we also performed a hand pollination experiment. 

We analyzed spatial associations among individuals, and the effects on reproductive performance and fruit removal of plant height, numbers of flowers and fruits, distance to forest edge, and neighboring flower and fruit density.

The data includes:

An R-script to perform analyses included in the article

      Arnell_2023_script.R

Spatial locations, reproductive status (vegetative / reproductive) and plant height of all individuals of D. mezereum in the local population in 2016-2017

       daphne_16.csv

       daphne_17.csv       

       daphne_18.csv

Data on number of flowers, number of fruits, fruit removal, plant height, flowers and fruits on neighboring reproductive individual within a 10 m radius of each reproductive individual and the distance to forest edges for each reproductive individual in 2016-2018

       reproductive_16.csv

       reproductive_17.csv       

       reproductive_18.csv

Data on flower gender (hermaphroditic/female) and fruit set for a subset of the local population in 2016-2018

       pollination_16.csv

       pollination_17.csv       

       pollination_18.csv

Results from the hand pollination experiment for a subset of the local population in 2019

       pollination_19.csv

Size of the study area and the length and location of forest edges (shapefiles) used in the analyses of spatial associations in R (each shapefile consists of six files with the extensions: .cpg, .dbf, .prj, .sbn, .sbx, .shp, .shx, and can be opened in R or any GIS prgram such as the open source software QGIS) 

        win.shp 

        lines.shp

Please refer to the README-files for further information about each data set / item.

STUDY SPECIES

Daphne mezereum is a deciduous shrub that can grow up to a height of 200 cm. In Sweden, D. mezereum is a relatively rare species found mainly on damp, humus-rich soils at forest edges or in deciduous forests previously managed as meadows or grazed by livestock. D. mezereum is one of the earliest species to flower in this region. Flowers are usually open from April to May, although during mild winters the onset of flowering can be as early as February.

LOCATION

This study was carried out on the peninsula of Väddö in the Stockholm archipelago in south eastern Sweden. The study area covers approximately 4.8 hectares, largely located in a forest. In the area, the forest is dominated by Picea abies, with deciduous species (e.g. Betula spp. and Populus tremula) occurring mostly in forest edges and in forest gaps. The soil consists of glacial till influenced by nearby occurring calcareous bedrock. The mean temperature of January is -3.5 C° and the mean temperature for June is -15.9 C°, with a mean annual precipitation of 595 mm.

FIELD SURVEY

During the years 2016-2018, we mapped all established individuals of D. mezereum of approximately 10 cm and higher within the study area. Individuals that were not detected during the first year were added continuously the following years of survey. In late April to early May we recorded the number of buds and flowers on each individual, as well as plant height. We recorded the number of fruits in June to August.

In 2019, to investigate if individuals were pollen limited, we surveyed 70 reproductive individuals. For each individual we recorded if flowers were hermaphroditic or female (anthers without pollen or lacking anthers). On individuals with more than one branch, we performed a cross-pollination experiment. One branch was pollinated and the rest was left un-treated as control. Pollen was taken from three or more randomly selected plants not included in the survey and applied with a small brush into each flower tube, on both hermaphroditic and female flowers. In April to May, we recorded number of flowers and in June the number of developing fruits.

POLLINATION AND FRUIT SET

To estimate the degree of pollen limitation, we compared fruit set between individuals having hermaphroditic flowers (producing pollen) and individuals having female flowers (producing no pollen), and assessed the effect of the pollen-addition treatment on fruit set.

SPATIAL ASSOCIATIONS

To describe the spatial structure of the population, we estimated spatial associations among individuals of D. mezereum based on the spatial locations of all individuals in the three survey years. Spatial associations were estimated using the inhomogeneous pair correlation function ginhom(r). Spatial associations among individuals with the same reproductive status (reproductive/vegetative) were also estimated separately. 

To describe the spatial structure of the population in relation to landscape structures, we fitted an inhomogeneous Thomas cluster process, with distance to forest edge (distance to edge) as a covariate. 

INVESTMENT IN REPRODUCTION, REPRODUCTIVE SUCCESS AND DISPERSAL IN RELATION TO DISPERSAL TRAITS AND DISTRIBUTION PATTERNS

We considered six predictor variables when investigating the effects of plant traits and distribution patterns on number of flowers (investment in reproduction), fruit set (relative reproductive success) and fruit removal. The predictors included three traits, height, number of flowers and number of fruits, and three measures of spatial locations, distance to forest edge, total number of flowers within a 10 m radius (neighboring flower density) and total number of fruits within a 10 m radius (neighboring fruit density). We chose a radius of 10 m as the analyses of spatial associations showed positive spatial associations up to 10 m.

The effect of three variables (height, distance to edge, neighboring flower density) were included in the model explaining the number of flowers produced by individuals of D. mezereum.

The effect of four variables (height, flowers, distance to edge, neighboring flower density) were included in the model explaining the fruit set of individuals of D. mezereum.

The effect of four variables (height, fruits, distance to edge, neighboring fruit density) were included in the model explaining the fruit removal in individuals of D. mezereum.

Please consult the original article as well as the R-script "Arnell_2023_script.R" for details on how the data was analyzed. 

Please contact Matilda Arnell (matilda.arnell@su.se) for information or collaboration. 

Please cite the original article when using these data (DOI: XXXXX).</description>
      <pubDate>Thu, 22 Jun 2023 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-21311154</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-21311154</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Matilda Arnell</dc:creator>
    </item>
    <item>
      <title>Data: Grazing livestock increases both vegetation and seed bank diversity in remnant and restored grasslands</title>
      <description>Location. Stockholm archipelagoData: Species list of the vegetation and seed bank with presence/absence</description>
      <pubDate>Wed, 01 Mar 2023 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-12962963</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-12962963</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Rozália Kapás</dc:creator>
      <dc:creator>Jan Plue</dc:creator>
      <dc:creator>Adam Kimberley</dc:creator>
      <dc:creator>Sara Cousins</dc:creator>
    </item>
    <item>
      <title>Data from: Landscape-scale range filling and dispersal limitation of woody plants</title>
      <description>This is the data from the article 

Landscape-scale range filling and dispersal limitation of woody plants (DOI: 10.1111/jbi.14485)

Matilda Arnell and Ove Eriksson

RANGE FILLING ESTIMATES

We estimated landscape-scale range filling for 64 species, each representing a different genera of woody plants, from two different dispersal systems:vertebrate dispersal and abiotic dispersal (mainly wind dispersed). 

Landscape-scale range filling was estimated as the proportion realized range within the potential range, at a 1km2 resolution.

We estimated potential ranges using species distribution models (SDMs) in continuous suitability scores (Seliger et al. 2020). This method avoids loss of information by not converting the SDM outputs into presence/absence using an arbitrary threshold of suitability. 

Realized ranges were estimated from presence records, restricting the estimations to areas with high sampling efforts: low ignorance areas (Ruete 2015), in order to increase the likelihood that absences represented true absences.  

Regional range filling was estimated for a 5000 pixel subset of the low ignorance areas. The aditional low ignorance pixels and accompanying occurence datat was used when trining the SDMs.

Please consult to the original article as well as the R-script "range filling analyses_Arnell_Eriksson_2022.R" for details on regional range filling estimates. 

LOCATION

We estimated regional range filling in the nemoral and boreo-nemoral vegetation zones in Sweden. The species distribution models providing the estimated suatability scores were trained with ocurrence data, climate and land-use data from all of Sweden.

PHYLOGENETIC REGRESSION

We thested the effect of dispersal system and habitat affinities on landscape-scale range filling using phylogenetic regressions. Phylogenetic information was obtained from Zanne et al. (2014). 

Please consult the original article as well as the R-script "PGLS models_Arnell_Eriksson_2022.R" for details on regional range filling estimates.  

HABITAT AFFINITIES

Plant indicator values (Tyler et al. 2021) used to assess the effect of habitat affinities:

Light indicator value

Moisture indicator value

Please contact Matilda Arnell (matilda.arnell@su.se) for information or collaboration. 

Please cite also the original article when using these data (DOI: 10.1111/jbi.14485).

REFERENCES

Ruete, A. (2015). Displaying bias in sampling effort of data accessed from biodiversity databases using ignorance maps. Biodiversity Data Journal, 3, e5361. https://doi.org/10.3897/BDJ.3.e5361

Seliger, B. J., McGill, B. J., Svenning, J., &amp; Gill, J. L. (2020). Widespread underfilling of the potential ranges of North American trees. Journal of Biogeography, 48(2), 359–371. https://doi.org/10.1111/jbi.14001

Tyler, T., Herbertsson, L., Olofsson, J., &amp; Olsson, P. A. (2021). Ecological indicator and traits values for Swedish vascular plants. Ecological Indicators, 120, 106923. https://doi.org/10.1016/j.ecolind.2020.106923

Zanne, A. E., Tank, D. C., Cornwell, W. K. et al. (2014). Three keys to the radiation of angiosperms into freezing environments. Nature, 506(7486), 89–92. https://doi.org/10.1038/nature12872</description>
      <pubDate>Mon, 10 Oct 2022 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-14784924</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-14784924</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Matilda Arnell</dc:creator>
      <dc:creator>Ove Eriksson</dc:creator>
    </item>
    <item>
      <title>COI reference sequences from BOLD DB</title>
      <description>This item contains COI (mitochondrial cytochrome oxidase subunit I) sequences collected from the BOLD database. The dataset is based on the BOLD Data Package from 15 May 2026.

The fasta file coidb.clustered.fasta.gz represents a non-redundant set of filtered sequences (clustered at 100% identity, see Methods) with record ids that can be queried in the Public Data Portal (https://portal.boldsystems.org/) . Each fasta header also contains the BIN ID assigned to the record (with the exception of prokaryotic records which instead have process ids as BIN IDs).

The taxonomic information for all filtered records is given in the tab-separated file coidb.info.tsv.gz.


Files compatible with specific tools for taxonomic assignments are found under the dada2/, sintax/, and qiime2/ folders.

MethodsThis dataset was generated with the coidb package (v0.7.0).

Briefly, records from the BOLD Data Package are filtered to:

- keep only records assigned a proper BOLD BIN (e.g. 'BOLD:AAA0008'), as well as records assigned to Bacteria or Archaea
- keep only records with marker_code 'COI-5P'

- remove records shorter than 500 bp
- remove records containing non-standard DNA characters



Remaining sequences are then clustered at 100% identity separately for each BOLD BIN using vsearch (v2.30.4, Rognes et al. 2016) (records without BOLD BINs that are assigned to Bacteria/Archaea are not clustered).

The taxonomic information for records is processed to handle missing data and non-unique parent lineages. A consensus taxonomy for each BOLD BIN is calculated by taking into account the taxonomic information given for records assigned to each BIN. This is done in two ways:

- the `inclNA` method calculates a consensus based on all taxonomic labels, even the ones with missing data
- the `exclNA` method excludes taxonomic labels with missing data when calculating the consensus



Because these methods have their pros and cons (in short exclNA resolves more species but inclNA is more conservative) both versions of downstream files are available in this item and it is up to the user to decide which one to use.

In addition, all unique species names were matched to the Catalogue of Life (https://www.gbif.org/dataset/7ddf754f-d193-4cc9-b351-99906754a03b)  checklist using the pygbif (https://pygbif.readthedocs.io/en/latest/index.html)  package (v0.6.6). Only records assigned to species that could be matched exactly and without ambiguity were kept and used to form the 'gbif' version of the database.

Description of filescoidb/coidb.clustered.fasta.gz
This file contains nucleotide sequences of all filtered records, clustered at 100% identity within each BOLD BIN. The fasta headers have the format:
&gt;{processid} bin_uri:{BOLD BIN}

where '{processid}' corresponds to the record identifier chosen as the cluster centroid and '{BOLD BIN}' shows which BOLD BIN the record belongs to.

coidb/coidb.info.tsv.gz
This file contains taxonomic information (including BOLD BIN where applicable) as well as nucleotide sequences for all filtered records.

gbif/gbif.info.tsv.gz
This file shows the result of matching species names from BOLD to the Catalogue of Life. The first column contains the species name from BOLD and subsequent columns show the matched taxonomic labels for ranks from kingdom -&gt; species. Unmatched species names have 'unassigned' as taxonomic labels.

consolidated/consolidated.tsv.gz
This file shows the complete information for each record remaining after matching species names to Catalogue of Life. It has the same format as the coidb/coidb.info.tsv.gz file.

consensus_taxonomy/coidb.exclNA.tsv.gz
consensus_taxonomy/coidb.inclNA.tsv.gz
consensus_taxonomy/gbif.exclNA.tsv.gz
consensus_taxonomy/gbif.inclNA.tsv.gz
These files contain the consensus taxonomy for BOLD BINs generated as described under Methods above. The files with the 'gbif' prefix contain only information for records assigned to species matched to the Catalogue of Life.

Tool-specific filesDADA2
The dada2/ folder contains fasta files that are compatible with the DADA2 assignTaxonomy and addSpecies functions. See more information at https://benjjneb.github.io/dada2/assign.html.

The files wtih 'toGenus' and 'toSpecies' in their names have taxonomic information down to the genus and species level, respectively. The files with 'addSpecies' contain only the species name and should be used with the 'addSpecies' function.

SINTAX
The sintax/ folder contains fasta files that are compatible with taxonomic assignments using the SINTAX algorithm as implemented in `vsearch`. See more information in the vsearch manual.

QIIME2

The qiime2/ folder contains info files that can be imported with QIIME2. For more information, see the README file at https://github.com/insect-biome-atlas/coidb.

Other fileslogs/fix_nonunique.coidb.log
logs/fix_nonunique.gbif.log
These files show how taxa with non-unique parent lineages were modified during database creation.

stats/general_stats.tsv
This file show statistics on the different databases in this upload. The columns are:

- type: the type of database (e.g., coidb.exclNA, gbif.exclNA etc)
- total_seqs: total number of sequences remaining after clustering
- total_bins: total number of unique BOLD BINs
- {mean,median,min,max}_seqs_per_bin: statistics on number of sequences per BOLD BIN (after clustering)
- total_non-bins: number of unique non-BOLD BINs. This typically represents prokaryotic records which are not assigned a BOLD BIN
- total_species: total number of unique species (also includes unresolved/ambiguous species names)
- total_bin_species: total number of unique species for sequences assigned a BOLD BIN
- total_nonbin_species: total number of unique species for sequences NOT assigned a BOLD BIN
- ambiguous_species: number of unique species with ambiguous taxonomic assignment (suffixed with "_X"). These are records with missing taxonomic information at species level.
- seqs_in_ambiguous_species: total number of sequences with ambiguous species names
- ambiguous_bin_species: same as ambiguous_species but only for BOLD BINs
- seqs_in_ambiguous_bin_species: same as seqs_in_ambiguous_species but only for BOLD BINs
- unresolved_species: total number of unique unresolved species (species names prefixed with 'unresolved.')
- seqs_in_unresolved_species: total number of sequences with unresolved species labels
- unresolved_bin_species: same as unresolved_species but only for BOLD BINs
- seqs_in_unresolved_bin_species: same as seqs_in_unresolved_species but only for BOLD BINs
- unresolved_ambiguous_species: total number of unresolved AND ambiguous species (species names prefixed with 'unresolved.' and suffixed with "_X").
- seqs_in_unresolved_ambiguous_species: total number of sequences with unresolved AND ambiguous species labels
- unresolved_ambiguous_bin_species: same as unresolved_ambiguous_species but only for BOLD BINs
seqs_in_unresolved_ambiguous_bin_species: same as seqs_in_unresolved_ambiguous_species but only for BOLD BINs


stats/taxa_stats.tsv
This file shows number of BOLD BINs and sequences for different taxa in each database. The file is in 'long-format' with columns:



- taxa: taxonomic name
- n_bins: number of unique BOLD BINs
- n_seqs: number of sequences (after clustering)
- rank: taxonomic rank (only kingdom and phylum are shown)
- name: database name
shasum.txt
This file contains checksums and can be used to verify file integrity by running

shasum -c shasum.txt</description>
      <pubDate>Thu, 15 Sep 2022 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-20514192</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-20514192</guid>
      <dc:publisher>Naturhistoriska riksmuseet</dc:publisher>
      <dc:creator>John Sundh</dc:creator>
    </item>
    <item>
      <title>Climate drives among-year variation in natural selection on flowering time</title>
      <description>data_indiv: This data set includes 22 years of data on flowering phenology and fitness of the perennial herb Lathyrus
vernus , as well as climatic data from the same years (see below). Plant data was collected in a population of L.
vernus in a deciduous forest in the
Tullgarn area, SE Sweden (58.9496 N, 17.6097 E), during 1987–1996
and 2006–2017. In
1987, all flowering individuals in an area of 825 m²
were permanently marked and surveyed in each year to 1996. New
flowering individuals in the plot were included in the study in each
year. No recordings were made 1997–2005. In 2006, a new set of
individuals in an area of 162 m²
within the same population were marked, and surveyed in the same way
as the initially marked individuals to 2017. In total, we recorded
2411 flowering events, and followed 606 individuals 1987–1996, and
228 individuals 2006–2017. 
climate_1961_2017: This data set includes climatic data from nearby stations for the period 1961-2017. Weather
data for March, April and May 1961–2017 was obtained from the
Swedish Meteorological and Hydrological Institute (www.smhi.se (http://www.smhi.se/) ). Daily mean, minimum and maximum temperature values were averaged from
two meteorological stations: Oxelösund (58.6777 N, 17.1223 E, 41 km
from the study population) and Södertälje (59.2142 N, 17.6289 E, 29
km from the study population). Daily
precipitation values were obtained from a third station located in
Åda (58.9279 N, 17.5358 E, 5 km from the study population). We calculated 12 variables from weather
data: monthly averages of daily minimum, mean and maximum
temperatures, and monthly sums of precipitation, for March, April and
May in the year of flowering. 
See the article for further information.</description>
      <pubDate>Wed, 05 Feb 2020 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-11576505</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-11576505</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Johan Ehrlén</dc:creator>
      <dc:creator>Alicia Valdés</dc:creator>
    </item>
  </channel>
</rss>