<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <atom:link rel="self" type="application/rss+xml" href="https://researchdata.se/sv/catalogue/search.rss?freeKeyword=Invertebrates"/>
    <link>https://researchdata.se/sv/catalogue</link>
    <title>Researchdata.se</title>
    <description>Search results</description>
    <language>sv</language>
    <item>
      <title>Tick abundance from Grimsö Research Area, 2014-05-01–2025-09-30</title>
      <description>Spring and Autumn tick count on individual rodents captured during the small rodent survey
Grimsö Wildlife Research Station (2026). Tick abundance from Grimsö Research Area, 2014-05-01–2025-09-30 [Data set]. Swedish Infrastructure for Ecosystem Science (SITES). https://hdl.handle.net/11676.1/QrnCQm8G4oohZiUTeJrctZFd</description>
      <pubDate>Thu, 22 Jan 2026 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/sites-qrncqm8g4oohziutejrctzfd</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/sites-qrncqm8g4oohziutejrctzfd</guid>
      <dc:publisher>Sveriges lantbruksuniversitet</dc:publisher>
    </item>
    <item>
      <title>Supporting data tracks for: "Breaking insect genome records: sequencing of Stylops ater (Strepsiptera) reveals a minute, compact genome with a reduced set of genes".</title>
      <description>This data set contains supporting data tracks for the manuscript: "Breaking insect genome records: sequencing of Stylops ater (Strepsiptera) reveals a minute, compact genome with a reduced set of genes". Assembly and gene annotation are available on ENA under the umbrella project PRJEB71963. Here we publish the following additional resources:

- repeatmasker.gff - repeat track generated via: a repeat library was modelled with the RepeatModeler2 [1] v2.0.2a package. As repeats can be part of actual protein-coding genes, the candidate repeats modelled by RepeatModeler were vetted against our protein set (minus transposons) to exclude any nucleotide motif stemming from low-complexity coding sequences. From the repeat library, identification of repeat sequences present in the genome was performed using RepeatMasker (https://www.repeatmasker.org/)  v4.1.5 [2]
- repeatrunner.gff - repeat track generated via: RepeatRunner [3]. RepeatRunner is a program that integrates RepeatMasker with BLASTX allowing analysing highly divergent repeats and divergent portions of repeats and identifying divergent protein coding portions of retro-elements and retroviruses not detected by RepeatMasker.
- trna.gff - tRNA track - have been predicted through tRNAscan version 1.4 [4].
- rfam.gff - ncRNA track - As the main source of information we use the RNA family database Rfam (version 14.9) [5]. Rfam provides curated co-variance (CM) models, which can be used together with the Infernal [6] package to predict ncRNAs in genomic sequences. By default, the set of CM profiles is limited by us to only included broadly conserved, eukaryotic ncRNA families. /! In general, Rfam-derived ‘annotations’ should rather be seen as ‘predictions’. With the exception of some very well conserved ncRNA families, many of the resulting Rfam predictions need to be considered with some care.
- ST_1.gff3 - Transcriptome assembly of Illumina RNA-seq library ST_1 (ENA: SAMEA12922144, ERX11689259) assembled using our in-house pipeline transcript_assembly (https://github.com/NBISweden/pipelines-nextflow/tree/master/subworkflows/transcript_assembly)  [7]. To minimise gene fusions in this parasite genome the maximum intron length was reduced from 500000 to 20000 (hisat2 --max-intronlen 20000). Otherwise default parameter were used.
- ST_2.gff3 - Transcriptome assembly of Illumina RNA-seq library ST_2 (ENA: SAMEA12922144, ERX11689261) assembled using our in-house pipeline transcript_assembly (https://github.com/NBISweden/pipelines-nextflow/tree/master/subworkflows/transcript_assembly)  [7]. To minimise gene fusions in this parasite genome the maximum intron length was reduced from 500000 to 20000 (hisat2 --max-intronlen 20000). Otherwise default parameter were used.
- rc2_evidence_abinitio.gff - gene models created from second MAKER run (rc2), combining evidence (from run 1 or rc1) and ab initio predictors. Specifically, AUGUSTUS was used for the rc2_evidence_abinitio.gff
- rc2_evidence_genemark.gff - gene models created from second MAKER run (rc2), combining evidence (from run 1 or rc1) and ab initio predictors. Specifically, GeneMark was used for the rc2_evidence_genemark.gff
References:

[1] - Flynn JM, Hubley R, Goubert C, Rosen J, Clark AG, Feschotte C, Smit AF. (2020) RepeatModeler2 for automated genomic discovery of transposable element families. Proceedings of the National Academy of Sciences. 117 (17) 9451-9457. https://doi.org/10.1073/pnas.1921046117

[2] - Smit AFA, Hubley R, Green, P. (2013-2015) RepeatMasker Open-4.0.

[3] - Yandell Lab: https://www.yandell-lab.org/software/repeatrunner.html

[4] - Lowe TM, Eddy SR. (1997) tRNAscan-SE: A program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Research 25(5): 955–964. https://doi.org/10.1093/nar/25.5.955 (https://doi.org/10.1093/nar/25.5.955) 

[5] - Ioanna Kalvari, Eric P Nawrocki, Nancy Ontiveros-Palacios, Joanna Argasinska, Kevin Lamkiewicz, Manja Marz, Sam Griffiths-Jones, Claire Toffano-Nioche, Daniel Gautheret, Zasha Weinberg, Elena Rivas, Sean R Eddy, Robert D Finn, Alex Bateman, Anton I Petrov, Rfam 14: expanded coverage of metagenomic, viral and microRNA families, Nucleic Acids Research, Volume 49, Issue D1, 8 January 2021, Pages D192–D200, https://doi.org/10.1093/nar/gkaa1047

[6] - The recommended citation for using Infernal 1.1 is E. P. Nawrocki and S. R. Eddy, Infernal 1.1: 100-fold faster RNA homology searches (http://eddylab.org/publications.html#Nawrocki13c) , Bioinformatics 29:2933-2935 (2013).

[7] - Github: https://github.com/NBISweden/pipelines-nextflow/tree/master/subworkflows/transcript_assembly</description>
      <pubDate>Fri, 14 Nov 2025 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-30604043</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-30604043</guid>
      <dc:publisher>Naturhistoriska riksmuseet</dc:publisher>
      <dc:creator>Johannes Bergsten</dc:creator>
      <dc:creator>Martin Pippel</dc:creator>
      <dc:creator>Meri LÃ¤hteenaro</dc:creator>
      <dc:creator>Julia Heintz</dc:creator>
      <dc:creator>Genevieve Diedericks</dc:creator>
      <dc:creator>Henrique G. Leitão</dc:creator>
      <dc:creator>Carlos Leyton Rotella</dc:creator>
      <dc:creator>Mahesh Binzer-Panchal</dc:creator>
      <dc:creator>Christian Tellgren-Roth</dc:creator>
      <dc:creator>Mai-Britt Mosbech</dc:creator>
      <dc:creator>Hannes Svardal</dc:creator>
      <dc:creator>Alice Mouton</dc:creator>
      <dc:creator>Giulio Formenti</dc:creator>
      <dc:creator>Ann M. Mc Cartney</dc:creator>
      <dc:creator>Henrik Lantz</dc:creator>
      <dc:creator>Olga Vinnere Pettersson</dc:creator>
    </item>
    <item>
      <title>Supplemental data from the genome assembly and annotation of the Clouded Apollo Butterfly (Parnassius mnemosyne)</title>
      <description>This dataset contains supplementary data from the genome sequencing of the Clouded Apollo Butterfly (Parnassius mnemosyne), published in:

Höglund, J., Dias, G., Olsen, R. A., Soares, A., Bunikis, I., Talla, V., &amp; Backström, N. (2024). A Chromosome-Level Genome Assembly and Annotation for the Clouded Apollo Butterfly (Parnassius mnemosyne): A Species of Global Conservation Concern. Genome Biology and Evolution, 16(2), evae031. https://doi.org/10.1093/gbe/evae031

Previous data from the project has been deposited at the European Nucleotide Archive (ENA) in the umbrella project PRJEB76269 (https://www.ebi.ac.uk/ena/browser/view/PRJEB76269) .

The data contained in this archive at SciLifeLab Data Repository describe the genome assembly (ENA accession: GCA_963668995.1 (https://www.ebi.ac.uk/ena/browser/view/GCA_963668995.1) ), and the mitochondrial genome assembly (ENA accession: OZ075093.1 (https://www.ebi.ac.uk/ena/browser/view/OZ075093.1) ).

Below follows a brief description of each file. The information on the methods used to generate the files was adapted from Höglund et al. 2024.

- pmne_functional_edit1.gff.gz
contains the functional annotation (protein coding genes) of the primary genome assembly (GCA_963668995.1 (https://www.ebi.ac.uk/ena/browser/view/GCA_963668995.1) ). This is the original file that was submitted to ENA. A derived version of the file is available from NCBI; the NCBI version was generated from the EMBL records of each annotated gene and differs in that it for instance use a different naming scheme for the seqid column and the locus tags. The NCBI version is available at this link (https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/963/668/995/GCA_963668995.1_Parnassius_mnemosyne_n_2023_11/GCA_963668995.1_Parnassius_mnemosyne_n_2023_11_genomic.gff.gz) .

The genes were predicted using BRAKER (v3.03), GALBA (v1.0.6), and GeneMarkS-T (v5.1). The resulting gene models were combined and filtered using TSEBRA (version: long_reads branch commit 1f2614). The combined gene model was functionally annotated by the NBIS nextflow pipeline v2.0.0 (https://github.com/NBISweden).

- pmne_Illumina_RNAseq_StringTie_sorted-transcripts_match.gff.gz
contains a transcript assembly of the Illumina RNAseq reads (ENA accession: ERX11559451 (https://www.ebi.ac.uk/ena/browser/view/ERX11559451) ). The reads were aligned to the genome with HiSat2 (v2.1.0) and then assembled with StringTie (v2.2.1).

- pmne_mtdna.gff.gz
contains the functional annotation of the mitochondrial genome assembly (ENA accession: OZ075093.1 (https://www.ebi.ac.uk/ena/browser/view/OZ075093.1) ). This is the original file that was submitted to ENA. The annotation was generated using MitoFinder (v1.4.1).

- pmne_ncRNAs.gff.gz
contains the annotation of putative non-coding RNA (ncRNA) genes. The prediction was done with Infernal (v1.1.4) and the Rfam (v14.1) covariance models.

- pmne_tRNAs_and_pseudogenes.gff.gz
contains the annotation of putative tRNA genes and pseudogenes. The prediction was done with tRNAscan-SE (v2.0.12).

- pmne_PacBio_isoseq.sorted.bam
contains the PacBio IsoSeq transcripts (ENA accession: ERX11559436 (https://www.ebi.ac.uk/ena/browser/view/ERX11559436) ) aligned to the primary genome assembly.

- pmne_repeat_library.fa.gz
contains the nucleotide sequences of the prediced repeats in fasta format. The prediction was done with RepeatModeler2 (v2.0.2a).

Available variablesFor a description of the column headers of the files, please see the following links to the documentation of the different file formats.

The GFF3 format (.gff) is described here: https://github.com/The-Sequence-Ontology/Specifications/blob/master/gff3.md

The BAM format (.bam) is a compressed version of the SAM format, both of which are described here: https://samtools.github.io/hts-specs/SAMv1.pdf

The fasta (.fa) format is described here: https://www.ncbi.nlm.nih.gov/genbank/fastaformat/

ContactFor questions about this dataset, please contact:
jacob.hoglund@ebc.uu.se
niclas.backstrom@ebc.uu.se</description>
      <pubDate>Wed, 26 Jun 2024 00:00:00 GMT</pubDate>
      <link>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-25908748</link>
      <guid>https://researchdata.se/sv/catalogue/dataset/doi-10-17044-scilifelab-25908748</guid>
      <dc:publisher>Uppsala universitet</dc:publisher>
      <dc:creator>Jacob Höglund</dc:creator>
      <dc:creator>Guilherme Dias</dc:creator>
      <dc:creator>Remi-André Olsen</dc:creator>
      <dc:creator>André Soares</dc:creator>
      <dc:creator>Ignas Bunikis</dc:creator>
      <dc:creator>Venkat Talla</dc:creator>
      <dc:creator>Niclas Backström</dc:creator>
    </item>
  </channel>
</rss>