Matches in Nanopublications for { ?s ?p ?o <https://w3id.org/np/RAuZdwi5v34TP1MWsdSxlqcFFrNgAsnNfPkF_bF-PjuGA/assertion>. }
- 0000-0003-2388-0744 type Agent assertion.
- Workflow-RO-Crate type CreativeWork assertion.
- about.workflowhub.eu type Organization assertion.
- bts480 type ComputerLanguage assertion.
- enrichment_service-account-enrichment type Agent assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df type Dataset assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df type ResearchObject assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df type LiveRO assertion.
- 17b98cad-68c7-4991-acb1-b922f0c3d44f type Dataset assertion.
- 17b98cad-68c7-4991-acb1-b922f0c3d44f type Folder assertion.
- 28d4cc6a-cc26-4629-865c-951dcec04a63 type Dataset assertion.
- 28d4cc6a-cc26-4629-865c-951dcec04a63 type Folder assertion.
- 72f47650-7679-4648-bb43-d28fa77665c4 type Dataset assertion.
- 72f47650-7679-4648-bb43-d28fa77665c4 type Folder assertion.
- ef0d1b26-8b1c-4a45-9758-52610cb8e3a8 type Dataset assertion.
- ef0d1b26-8b1c-4a45-9758-52610cb8e3a8 type Folder assertion.
- 043ddb69-1a1c-4eb4-8785-9e3803fa377c type CreativeWork assertion.
- 043ddb69-1a1c-4eb4-8785-9e3803fa377c type MediaObject assertion.
- 043ddb69-1a1c-4eb4-8785-9e3803fa377c type Resource assertion.
- 0cdb95da-31b8-40ad-a2f9-32154e335db2 type MediaObject assertion.
- 0cdb95da-31b8-40ad-a2f9-32154e335db2 type Resource assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 type SoftwareSourceCode assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 type MediaObject assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 type Resource assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 type ComputationalWorkflow assertion.
- 1dd2de01-f16f-4704-853a-116e2de3ff65 type MediaObject assertion.
- 1dd2de01-f16f-4704-853a-116e2de3ff65 type Resource assertion.
- 21e37e3d-1372-4acd-9ce7-b1a005e9a41a type MediaObject assertion.
- 21e37e3d-1372-4acd-9ce7-b1a005e9a41a type Resource assertion.
- 2dd40f39-d265-4afe-ab50-5238e7bd6b16 type MediaObject assertion.
- 2dd40f39-d265-4afe-ab50-5238e7bd6b16 type Resource assertion.
- 52497c90-bae5-4fbd-aca5-b5175ff7a4fa type MediaObject assertion.
- 52497c90-bae5-4fbd-aca5-b5175ff7a4fa type Resource assertion.
- 5e018f3a-bc34-48db-b72e-f85214a13a9d type MediaObject assertion.
- 5e018f3a-bc34-48db-b72e-f85214a13a9d type Resource assertion.
- 6420e76a-1e0e-42a0-9b9e-e355de2a88a5 type MediaObject assertion.
- 6420e76a-1e0e-42a0-9b9e-e355de2a88a5 type Resource assertion.
- 64cb8be0-e907-41a0-8f6b-dbf642fe4372 type MediaObject assertion.
- 64cb8be0-e907-41a0-8f6b-dbf642fe4372 type Resource assertion.
- 69f9d60c-925c-4b44-9b87-fc35c96eed2f type MediaObject assertion.
- 69f9d60c-925c-4b44-9b87-fc35c96eed2f type Resource assertion.
- 6aadc530-3b55-44e0-8c8d-8762952df493 type MediaObject assertion.
- 6aadc530-3b55-44e0-8c8d-8762952df493 type Resource assertion.
- 70cec592-1fbe-4b53-a8ef-8dd3522246d6 type MediaObject assertion.
- 70cec592-1fbe-4b53-a8ef-8dd3522246d6 type Resource assertion.
- 7d430516-d8cb-4bc0-a4d9-e103a63e3478 type MediaObject assertion.
- 7d430516-d8cb-4bc0-a4d9-e103a63e3478 type Resource assertion.
- 89e1f03d-e2e9-46f1-8386-8cdeff35386c type MediaObject assertion.
- 89e1f03d-e2e9-46f1-8386-8cdeff35386c type Resource assertion.
- b2864dbb-5ce6-4ad9-8dcb-27a0325baeb1 type MediaObject assertion.
- b2864dbb-5ce6-4ad9-8dcb-27a0325baeb1 type Resource assertion.
- d6d97bc9-1631-4347-920b-dec30f6aebea type MediaObject assertion.
- d6d97bc9-1631-4347-920b-dec30f6aebea type Resource assertion.
- ro-crate-metadata.json type CreativeWork assertion.
- 0000-0003-4771-6113 type Person assertion.
- 554 type Person assertion.
- 191 type Organization assertion.
- 191 type Project assertion.
- 0000-0003-2388-0744 name "Małgorzata Wolniewicz" assertion.
- Workflow-RO-Crate name "Workflow RO-Crate Profile" assertion.
- about.workflowhub.eu name "WorkflowHub" assertion.
- bts480 name "Snakemake" assertion.
- enrichment_service-account-enrichment name "service-account-enrichment" assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df name "Research Object Crate for ERGA Protein-coding gene annotation workflow" assertion.
- 17b98cad-68c7-4991-acb1-b922f0c3d44f name "files" assertion.
- 28d4cc6a-cc26-4629-865c-951dcec04a63 name "rules" assertion.
- 72f47650-7679-4648-bb43-d28fa77665c4 name "envs" assertion.
- ef0d1b26-8b1c-4a45-9758-52610cb8e3a8 name "config" assertion.
- 043ddb69-1a1c-4eb4-8785-9e3803fa377c name "ro-crate-preview.html" assertion.
- 0cdb95da-31b8-40ad-a2f9-32154e335db2 name "hisat.yaml" assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 name "ERGA Protein-coding gene annotation workflow" assertion.
- 1dd2de01-f16f-4704-853a-116e2de3ff65 name "fastqc.yaml" assertion.
- 21e37e3d-1372-4acd-9ce7-b1a005e9a41a name "protexcluder.yaml" assertion.
- 2dd40f39-d265-4afe-ab50-5238e7bd6b16 name "uniprot_sprot.fasta.gz" assertion.
- 52497c90-bae5-4fbd-aca5-b5175ff7a4fa name "1_MaskRepeat.smk" assertion.
- 5e018f3a-bc34-48db-b72e-f85214a13a9d name "tePSI.yaml" assertion.
- 6420e76a-1e0e-42a0-9b9e-e355de2a88a5 name "trimm.yaml" assertion.
- 64cb8be0-e907-41a0-8f6b-dbf642fe4372 name "blast.yaml" assertion.
- 69f9d60c-925c-4b44-9b87-fc35c96eed2f name "agat.yaml" assertion.
- 6aadc530-3b55-44e0-8c8d-8762952df493 name "samtools.yaml" assertion.
- 70cec592-1fbe-4b53-a8ef-8dd3522246d6 name "gaas.yaml" assertion.
- 7d430516-d8cb-4bc0-a4d9-e103a63e3478 name "2_alignRNA.smk" assertion.
- 89e1f03d-e2e9-46f1-8386-8cdeff35386c name "README.md" assertion.
- b2864dbb-5ce6-4ad9-8dcb-27a0325baeb1 name "3_braker.smk" assertion.
- d6d97bc9-1631-4347-920b-dec30f6aebea name "config.yaml" assertion.
- 0000-0003-4771-6113 name "Sagane Joye-Dind" assertion.
- 554 name "Sagane Joye-Dind" assertion.
- 191 name "ERGA Annotation" assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 conformsTo "https://bioschemas.org/profiles/ComputationalWorkflow/1.0-RELEASE/" assertion.
- ro-crate-metadata.json conformsTo 1.1 assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df importedBy 0000-0003-2388-0744 assertion.
- about.workflowhub.eu url "https://about.workflowhub.eu/" assertion.
- bts480 url "https://snakemake.readthedocs.io/" assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df url "https://workflowhub.eu/workflows/569/ro_crate?version=1" assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 url "https://github.com/ERGA-consortium/pipelines/tree/main/annotation/snakemake/Snakefile" assertion.
- 043ddb69-1a1c-4eb4-8785-9e3803fa377c about "https://ror.org/https://doi.org/10.48546/workflowhub.workflow.569.1" assertion.
- ro-crate-metadata.json about 7019724e-b5a0-4f7e-a7d6-a1baacac85df assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df description "# ERGA Protein-coding gene annotation workflow. Adapted from the work of Sagane Joye: https://github.com/sdind/genome_annotation_workflow ## Prerequisites The following programs are required to run the workflow and the listed version were tested. It should be noted that older versions of snakemake are not compatible with newer versions of singularity as is noted here: [https://github.com/nextflow-io/nextflow/issues/1659](https://github.com/nextflow-io/nextflow/issues/1659). `conda v 23.7.3` `singularity v 3.7.3` `snakemake v 7.32.3` You will also need to acquire a licence key for Genemark and place this in your home directory with name `~/.gm_key` The key file can be obtained from the following location, where the licence should be read and agreed to: http://topaz.gatech.edu/GeneMark/license_download.cgi ## Workflow The pipeline is based on braker3 and was tested on the following dataset from Drosophila melanogaster: [https://doi.org/10.5281/zenodo.8013373](https://doi.org/10.5281/zenodo.8013373) ### Input data - Reference genome in fasta format - RNAseq data in paired-end zipped fastq format - uniprot fasta sequences in zipped fasta format ### Pipeline steps - **Repeat Model and Mask** Run RepeatModeler using the genome as input, filter any repeats also annotated as protein sequences in the uniprot database and use this filtered libray to mask the genome with RepeatMasker - **Map RNAseq data** Trim any remaining adapter sequences and map the trimmed reads to the input genome - **Run gene prediction software** Use the mapped RNAseq reads and the uniprot sequences to create hints for gene prediction using Braker3 on the masked genome - **Evaluate annotation** Run BUSCO to evaluate the completeness of the annotation produced ### Output data - FastQC reports for input RNAseq data before and after adapter trimming - RepeatMasker report containing quantity of masked sequence and distribution among TE families - Protein-coding gene annotation file in gff3 format - BUSCO summary of annotated sequences ## Setup Your data should be placed in the `data` folder, with the reference genome in the folder `data/ref` and the transcript data in the foler `data/rnaseq`. The config file requires the following to be given: ``` asm: 'absolute path to reference fasta' snakemake_dir_path: 'path to snakemake working directory' name: 'name for project, e.g. mHomSap1' RNA_dir: 'absolute path to rnaseq directory' busco_phylum: 'busco database to use for evaluation e.g. mammalia_odb10' ```" assertion.
- 7019724e-b5a0-4f7e-a7d6-a1baacac85df description "# ERGA Protein-coding gene annotation workflow. Adapted from the work of Sagane Joye: https://github.com/sdind/genome_annotation_workflow ## Prerequisites The following programs are required to run the workflow and the listed version were tested. It should be noted that older versions of snakemake are not compatible with newer versions of singularity as is noted here: [https://github.com/nextflow-io/nextflow/issues/1659](https://github.com/nextflow-io/nextflow/issues/1659). `conda v 23.7.3` `singularity v 3.7.3` `snakemake v 7.32.3` You will also need to acquire a licence key for Genemark and place this in your home directory with name `~/.gm_key` The key file can be obtained from the following location, where the licence should be read and agreed to: http://topaz.gatech.edu/GeneMark/license_download.cgi ## Workflow The pipeline is based on braker3 and was tested on the following dataset from Drosophila melanogaster: [https://doi.org/10.5281/zenodo.8013373](https://doi.org/10.5281/zenodo.8013373) ### Input data - Reference genome in fasta format - RNAseq data in paired-end zipped fastq format - uniprot fasta sequences in zipped fasta format ### Pipeline steps - **Repeat Model and Mask** Run RepeatModeler using the genome as input, filter any repeats also annotated as protein sequences in the uniprot database and use this filtered libray to mask the genome with RepeatMasker - **Map RNAseq data** Trim any remaining adapter sequences and map the trimmed reads to the input genome - **Run gene prediction software** Use the mapped RNAseq reads and the uniprot sequences to create hints for gene prediction using Braker3 on the masked genome - **Evaluate annotation** Run BUSCO to evaluate the completeness of the annotation produced ### Output data - FastQC reports for input RNAseq data before and after adapter trimming - RepeatMasker report containing quantity of masked sequence and distribution among TE families - Protein-coding gene annotation file in gff3 format - BUSCO summary of annotated sequences ## Setup Your data should be placed in the `data` folder, with the reference genome in the folder `data/ref` and the transcript data in the foler `data/rnaseq`. The config file requires the following to be given: ``` asm: 'absolute path to reference fasta' snakemake_dir_path: 'path to snakemake working directory' name: 'name for project, e.g. mHomSap1' RNA_dir: 'absolute path to rnaseq directory' busco_phylum: 'busco database to use for evaluation e.g. mammalia_odb10' ``` " assertion.
- 11f3e069-8d1d-48da-876f-52fd6d255223 description "# ERGA Protein-coding gene annotation workflow. Adapted from the work of Sagane Joye: https://github.com/sdind/genome_annotation_workflow ## Prerequisites The following programs are required to run the workflow and the listed version were tested. It should be noted that older versions of snakemake are not compatible with newer versions of singularity as is noted here: [https://github.com/nextflow-io/nextflow/issues/1659](https://github.com/nextflow-io/nextflow/issues/1659). `conda v 23.7.3` `singularity v 3.7.3` `snakemake v 7.32.3` You will also need to acquire a licence key for Genemark and place this in your home directory with name `~/.gm_key` The key file can be obtained from the following location, where the licence should be read and agreed to: http://topaz.gatech.edu/GeneMark/license_download.cgi ## Workflow The pipeline is based on braker3 and was tested on the following dataset from Drosophila melanogaster: [https://doi.org/10.5281/zenodo.8013373](https://doi.org/10.5281/zenodo.8013373) ### Input data - Reference genome in fasta format - RNAseq data in paired-end zipped fastq format - uniprot fasta sequences in zipped fasta format ### Pipeline steps - **Repeat Model and Mask** Run RepeatModeler using the genome as input, filter any repeats also annotated as protein sequences in the uniprot database and use this filtered libray to mask the genome with RepeatMasker - **Map RNAseq data** Trim any remaining adapter sequences and map the trimmed reads to the input genome - **Run gene prediction software** Use the mapped RNAseq reads and the uniprot sequences to create hints for gene prediction using Braker3 on the masked genome - **Evaluate annotation** Run BUSCO to evaluate the completeness of the annotation produced ### Output data - FastQC reports for input RNAseq data before and after adapter trimming - RepeatMasker report containing quantity of masked sequence and distribution among TE families - Protein-coding gene annotation file in gff3 format - BUSCO summary of annotated sequences ## Setup Your data should be placed in the `data` folder, with the reference genome in the folder `data/ref` and the transcript data in the foler `data/rnaseq`. The config file requires the following to be given: ``` asm: 'absolute path to reference fasta' snakemake_dir_path: 'path to snakemake working directory' name: 'name for project, e.g. mHomSap1' RNA_dir: 'absolute path to rnaseq directory' busco_phylum: 'busco database to use for evaluation e.g. mammalia_odb10' ```" assertion.