This page explains each input field on the job submission form. Click the Help link next to any field to jump directly to the relevant section below.
Upload the genome sequence of your organism as a multi-FASTA file.
Each sequence record must begin with a header line starting with
>, followed by the nucleotide sequence.
Vipsania predicts genes on all sequences (scaffolds, chromosomes, contigs)
present in the file.
.fa.gz or .fasta.gz: gzip-compressed FASTAThe maximum accepted upload size is 500 MB, and the genome may contain at most 300 Mbp. Larger genomes might not finish within the runtime limit of this server, because jobs run on CPU whenever no GPU is free; please run Vipsania locally for those.
Vipsania reads the soft masking of a genome, its lowercase nucleotides, as repeats. Good masking helps, but masking that is too aggressive or too sparse misleads the model. Upload your genome with the masking you trust, or without masking (all uppercase).
The job status page shows an MD5 checksum of the genome file exactly as it was received by the server. You can verify that your file was transferred without corruption by comparing this value against the MD5 you compute locally:
md5sum your_genome.fa.gz
If the two checksums match, the server processed the exact file you uploaded. This is particularly useful when sharing a job link with a collaborator — they can confirm which genome version was used.
A small example genome suitable for a test submission is available here:
example_genome.fa.gz
(8 scaffolds from the BRAKER example data set; use with the
Streptophyta clade model).
Once your job completes, the status page provides download links for the following files:
vipsania.gff3.gz — gene predictions in GFF3 formatcoding.fasta.gz — predicted coding sequences (CDS) in FASTA formatproteins.fasta.gz — predicted protein sequences in FASTA format
All files are gzip-compressed. Decompress them with gunzip
before loading them into downstream tools that do not accept compressed input:
gunzip vipsania.gff3.gz gunzip coding.fasta.gz gunzip proteins.fasta.gz
Many tools (genome browsers, BLAST, BUSCO, etc.) can read
.gz files directly — check the documentation of your tool
before decompressing.
Vipsania is trained per clade, on raw genome sequences only. Select the most specific clade that contains the organism whose genome you are annotating. Because Vipsania never learns from annotations, finding your species or a close relative among the training species of a model is not an argument against using that model.
| Clade | Suitable for | Model ID | Locus F1 on this server (pretrained) |
Locus F1 with finetuning (not offered here) |
| Alveolata | Ciliates, apicomplexans, dinoflagellates (Plasmodium, Tetrahymena, …) | sd2zcj7u |
0.582 | 0.613 |
| Amoebozoa | Amoebae and slime moulds (Dictyostelium, Entamoeba, …) | ezhpj2qm |
0.416 | 0.437 |
| Arthropoda (other than insects) | Crustaceans, arachnids, myriapods; for insects use Insecta | ihe1jk30 |
0.217 | 0.207 |
| Chlorophyta (green algae) | Green algae (Chlamydomonas, …) | faeijtmk |
0.558 | 0.579 |
| Cnidaria | Corals, sea anemones, jellyfish (Nematostella, Hydra, …) | b5vtieo0 |
0.406 | 0.453 |
| Discoba | Kinetoplastids, euglenids, heteroloboseans (Trypanosoma, Leishmania, Naegleria, …) | gcra9d9y |
0.696 | 0.696 |
| Echinodermata | Sea urchins, sea stars, sea cucumbers | sx9zjl7p |
0.440 | 0.444 |
| Fungi | Fungi (Saccharomyces, Aspergillus, …) | fh1kg88z |
0.645 | 0.660 |
| Insecta | Insects (Drosophila, Apis, …) | 58hsuobw |
0.474 | 0.537 |
| Nematoda | Roundworms (Caenorhabditis, …) | zspca2rb |
0.517 | 0.539 |
| Porifera | Sponges (Amphimedon, …) | 7bjiexcu |
0.416 | 0.480 |
| Rhodophyta (red algae) | Red algae (Porphyra, Cyanidioschyzon, …) | 6kmw3wme |
0.301 | 0.405 |
| Spiralia | Molluscs, annelids, flatworms | v5ej8oyt |
0.262 | 0.374 |
| Stramenopiles | Diatoms, brown algae, oomycetes (Phaeodactylum, Phytophthora, …) | r6p9z9jw |
0.413 | 0.426 |
| Streptophyta (land plants and relatives) | Land plants and charophyte algae (Arabidopsis, rice, …) | j9m0cdmk |
0.579 | 0.602 |
| Tunicata | Tunicates (Ciona, …) | hcehc7ff |
0.494 | 0.523 |
| Vertebrata | Fish, amphibians, reptiles, birds, mammals | etb1go6q |
0.327 | 0.406 |
| Other eukaryotes (none of the above) | Eukaryotes that belong to none of the clades listed above | cg6grhms |
0.416 | 0.436 |
The F1 columns are the average locus-level F1 over the test species of each clade, measured by the Vipsania authors against reference annotations (details). If your organism belongs to none of the clades, choose Other eukaryotes; that model was trained on exactly such species.
Since Vipsania is unsupervised, it can be trained further on the very genome
it is about to annotate (--finetune). The Vipsania authors
recommend this, and in their benchmark it improved the average locus F1 in
almost every clade (compare the two F1 columns above).
This web server does not finetune, and there is no option to enable it. Finetuning is a GPU training run per job, and we do not have the GPU capacity to offer that as a public service. All jobs are annotated with the pretrained clade model. If you need the additional accuracy that finetuning may provide, please run Vipsania on your own hardware; see the tutorial.
Provide a valid e-mail address so that we can send you:
Your e-mail address is deleted as soon as the result or error notification has been sent. If sending fails for any reason, it is deleted at the latest after 45 days (when all job files are removed anyway). It is never used for any purpose other than result notification. See the Data Privacy page for full details.
We do not verify e-mail addresses at submission time; please ensure the address is correct, or you will not receive your results.
By checking this box you confirm that the genome you are uploading is not a personalized human genome sequence — i.e. it does not contain genomic data derived from an identifiable individual.
Reference assemblies (e.g. GRCh38) and non-human genomes are not affected by this restriction. Uploading personalized human genomic data (e.g. a whole-genome sequencing result from a patient or research participant) is prohibited because this service does not provide the safeguards required for the processing of sensitive personal health data under EU GDPR.
Before submitting a job you must confirm that you have read and agree to our Data Privacy Protection Declaration.
The data processed during your job submission are:
No data are shared with third parties. The server is operated by the Bioinformatics Group at Universität Greifswald, Germany, and is subject to EU GDPR regulations.