<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <atom:link rel="self" type="application/rss+xml" href="https://researchdata.se/sv/catalogue/search.rss?freeKeyword=Biogeography+and+phylogeography"/>
    <link>https://researchdata.se/sv/catalogue</link>
    <title>Researchdata.se</title>
    <description>Search results</description>
    <language>sv</language>
    <item>
      <title>Phylogenomics of aquatic bacteria</title>
      <description>Intermediate data files obtained during the work on the manuscript "Phylogenomics of aquatic bacteria reveal molecular mechanisms behind the limits of their adaptation to salinity". The files published here were used at various stages of the analysis (or sum-up the stages) and should allow reproduction of the results as well as expanded investigation of the dataset. 

These files are:

ar_mags_info.txt  - information table for collected archeal MAGs. It contains names of the MAGs in format {data source 2-letter code}_{name of the MAG as in  ENA}, the biome of origin and taxonomic classification. For the brackish MAGs there is also annotation to the basin (Baltic/Caspian) of origin and additional metadata for the Baltic Sea MAGs.

bac_mags_info.txt - information table for collected bacterial MAGs. It contains names of the MAGs in format {data source 2-letter code}_{name of the MAG as in  ENA}, the biome of origin and taxonomic classification. For the brackish MAGs there is also annotation to the basin (Baltic/Caspian) of origin and additional metadata for the Baltic Sea MAGs.

CheckM_all_MAGs.csv - CheckM results for all the investigated MAGs (completness, contamination, strain heterogeneity).

ani_file.txt - average nucleotide indentity between all the pairs of investigated MAGs.

MAG-cluster-stats-interbiome-clusters.xlsx - Excel file with a table annotating MAGs to &gt;95% ANI clusters and the represtatives chosen for further analysis marked. Contains also sheets with just the representatives, clusters common between the brackish basins and between the biomes, as well as MSG_table.tsv imported into Excel spreadsheet. The first sheet also contains accession numbers for the bacterial MAGs used in this study. 

nozero.bifurc.bac.tree.nwk - phylogenetic tree of all the MAGs and GTDB reference genomes. Obtained using GTDB-tk.

pruned95.nozero.bifurc.bac.tree.nwk - the phylogenetic tree (nozero.bifurc.bac.tree.nwk) pruned to contain only one represtative for a biome from each &gt;95% ANI cluster. Does not contain GTDB reference genomes.

subsampled.pruned95.nozero.bifurc.bac.tree.nwk - the phylogenetic tree with &gt;95% ANI cluster respresntatives further randomly pruned the same number of freshwater and marine representatives.

timetree_evo_rate_100.nwk  -  the full phylogenetic tree (nozero.bifurc.bac.tree.nwk) with branch length adjusted to correspond to estimated times since divergence in mya [million years ago].

time_calibration.txt - constraint file used for estimating time since divergence, input for RelTime (MEGA11). Minimal estimates of time since host species diverged [mya], based on the fossil record, were used to set the constraints 

MSG_table.tsv - a table (tab-separated) with all the MAGs within identified monobiomic sister groups (MSGs), annotated to appropriate transition_ID, biome and transition type. Taxonomic classification and transition times and directions are also included.

make_MSG_table.R - R script used to make MSG_table.tsv.

assess_datetree.R - R scirpt used to find the cross-biome transitions on the time-adjusted phylogenetic tree and obtained the information about the estimated time since they occured.

transition_directions.R  - R script used to estimate the ancestral biome-states of MRCAs (most recent common ancestors) of MSG pair and thus infer the most probable transition directions.

all_MSG_ids.txt - a text file with names of all the representative MAGs within all the MSG pairs.

filter_MSGs.py - a Python script to extract the MAGs from within the MSGs (given all_MSG_ids.txt) from a folder containing a larger set of sequences.

Snakefile_proteins - Snakefile with a pipeline to go from nucleotide MAG sequences to pepstats statistics for inferred proteins. Includes proteome inference step using Prodigal (same procedure was used to infer amino-acid sequences for other purposes, including the taxonomic classification and reconstruction of the phylogenetic tree).

MSGs_whole_proteomes.py - a Python script to concatanate inferred proteomes into continous amino acid sequences (for amino acid usage statistics).

Snakefile_whole_proteome - Snakefile with a pipeline to obtain amino acid relative frequencies within proteomes. As an input takes proteomes in form of one continous sequence (MSGs_whole_proteomes.py output).

MSGs_pI_rel_freq_table.tsv  -  a table (tab separated) with relative frequencies of proteins with pIs (isoelectric points) within 0.5 pH wide bins.

aa_freqs_MSGs_list.json and assessed_aas.tsv - a json file with amino acid relative frequencies for each inferred proteome and a tab-separated list of IUPAC amino acid codes in order corresponding to values in the list.

aa_cat_freqs_MSGs_list.json and assessed_aa_cats.tsv - a json file with relative frequencies of amino acid categories for each inferred proteome and a tab-separated list names of the categories ordered accoridngly as in the .json file.

pI_aa_statistics.xlsx - statistics (p-values and differences sizes) for pairwise comparisons of inferred proteome properties and composition, i.e. i) relative frequencies of acidic, neutral and basic (isoelectric point (pI) categories) proteins ; ii) genome sizes as defined by number of inferred protein-coding genes; iii) amino acid relative frequencies; iv) relative frequencies of amino acids categories.

{transition type}.annotation.gz and MSG_ids_{transition type}_pairs.txt - annotation files (zipped) of inferred genes for random pairs of MAGs from each MSG pair, together with text file with MSG represntatives. Seperate pair of files for each transition type. Used for investigating coannotation.

ko_annot_full_everything.tsv - table with multilevel annotation of KEGG orthologs, adpoted from KEGG orthology website

ko_anno.rar - compressed table with numbers of genes annotated to respective KEGG orthologs in representative genomes from all the &gt;95% ANI bacterial clusters (used for gene gain/loss analysis).

iterate_rarefying.R - R script used to indentify the differentially present genes.

gain_loss_tables.xlsx - Sheets 1-3: results of MSG-based (phylogeny-aware) gene content analysis.  Tables with all the significant (FDR &lt; 0.1, shaded in orange) differentially present genes across pairs of MSGs. For FB and FM type transitions additional genes were added to the table to show at least the top 25 most significant genes regardless of the FDR values. Sheets 4-6: Biome(s) in which the differentially present KOs were found across the identified transitions (MSG pairs), i.e. the data presented in Fig. 6 in text form and annotated to more specific taxa and single transition events. Includes taxonomic annotation of the transitions and numbers of bacterial species in MSGs from respective biomes. Sheets 7-9: Fraction of cases in which gene A (row) was also annotated as gene B (column), based on {transition type}.annotation.gz files. Sheets 10-12: Results of phylogeny-unaware gene content analysis. Tables with all the significant (FDR &lt; 0.1) differentially present genes from an unpaired comparison of all bacterial species from each biome.</description>
      <pubDate>Mon, 17 Feb 2025 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-20732170</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-20732170</guid>
      <dc:publisher>Kungliga Tekniska högskolan</dc:publisher>
      <dc:creator>Krzysztof Jurdzinski</dc:creator>
      <dc:creator>Maliheh Mehrshad</dc:creator>
      <dc:creator>Stefan Bertilsson</dc:creator>
      <dc:creator>Anders Andersson</dc:creator>
    </item>
    <item>
      <title>Data from: Landscape-scale range filling and dispersal limitation of woody plants</title>
      <description>This is the data from the article 

Landscape-scale range filling and dispersal limitation of woody plants (DOI: 10.1111/jbi.14485)

Matilda Arnell and Ove Eriksson

RANGE FILLING ESTIMATES

We estimated landscape-scale range filling for 64 species, each representing a different genera of woody plants, from two different dispersal systems:vertebrate dispersal and abiotic dispersal (mainly wind dispersed). 

Landscape-scale range filling was estimated as the proportion realized range within the potential range, at a 1km2 resolution.

We estimated potential ranges using species distribution models (SDMs) in continuous suitability scores (Seliger et al. 2020). This method avoids loss of information by not converting the SDM outputs into presence/absence using an arbitrary threshold of suitability. 

Realized ranges were estimated from presence records, restricting the estimations to areas with high sampling efforts: low ignorance areas (Ruete 2015), in order to increase the likelihood that absences represented true absences.  

Regional range filling was estimated for a 5000 pixel subset of the low ignorance areas. The aditional low ignorance pixels and accompanying occurence datat was used when trining the SDMs.

Please consult to the original article as well as the R-script "range filling analyses_Arnell_Eriksson_2022.R" for details on regional range filling estimates. 

LOCATION

We estimated regional range filling in the nemoral and boreo-nemoral vegetation zones in Sweden. The species distribution models providing the estimated suatability scores were trained with ocurrence data, climate and land-use data from all of Sweden.

PHYLOGENETIC REGRESSION

We thested the effect of dispersal system and habitat affinities on landscape-scale range filling using phylogenetic regressions. Phylogenetic information was obtained from Zanne et al. (2014). 

Please consult the original article as well as the R-script "PGLS models_Arnell_Eriksson_2022.R" for details on regional range filling estimates.  

HABITAT AFFINITIES

Plant indicator values (Tyler et al. 2021) used to assess the effect of habitat affinities:

Light indicator value

Moisture indicator value

Please contact Matilda Arnell (matilda.arnell@su.se) for information or collaboration. 

Please cite also the original article when using these data (DOI: 10.1111/jbi.14485).

REFERENCES

Ruete, A. (2015). Displaying bias in sampling effort of data accessed from biodiversity databases using ignorance maps. Biodiversity Data Journal, 3, e5361. https://doi.org/10.3897/BDJ.3.e5361

Seliger, B. J., McGill, B. J., Svenning, J., &amp; Gill, J. L. (2020). Widespread underfilling of the potential ranges of North American trees. Journal of Biogeography, 48(2), 359–371. https://doi.org/10.1111/jbi.14001

Tyler, T., Herbertsson, L., Olofsson, J., &amp; Olsson, P. A. (2021). Ecological indicator and traits values for Swedish vascular plants. Ecological Indicators, 120, 106923. https://doi.org/10.1016/j.ecolind.2020.106923

Zanne, A. E., Tank, D. C., Cornwell, W. K. et al. (2014). Three keys to the radiation of angiosperms into freezing environments. Nature, 506(7486), 89–92. https://doi.org/10.1038/nature12872</description>
      <pubDate>Mon, 10 Oct 2022 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-14784924</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17045-sthlmuni-14784924</guid>
      <dc:publisher>Stockholms universitet</dc:publisher>
      <dc:creator>Matilda Arnell</dc:creator>
      <dc:creator>Ove Eriksson</dc:creator>
    </item>
  </channel>
</rss>