Help

This page explains each input field on the job submission form. Click the Help link next to any field to jump directly to the relevant section below.

No finetuning on this server. Vipsania can be finetuned on the genome it is about to annotate (vipsania annotate --finetune), and its authors recommend doing so. This web server does not offer finetuning, because we do not have the GPU capacity for it: every job is annotated with the pretrained clade model as it is. Finetuning may make the annotation somewhat more accurate. If you need that, please run Vipsania locally with --finetune; a GPU is strongly recommended for finetuning.

Genome FASTA file

Upload the genome sequence of your organism as a multi-FASTA file. Each sequence record must begin with a header line starting with >, followed by the nucleotide sequence. Vipsania predicts genes on all sequences (scaffolds, chromosomes, contigs) present in the file.

Accepted formats

  • .fa.gz or .fasta.gz: gzip-compressed FASTA

Size limit

The maximum accepted upload size is 500 MB, and the genome may contain at most 300 Mbp. Larger genomes might not finish within the runtime limit of this server, because jobs run on CPU whenever no GPU is free; please run Vipsania locally for those.

Requirements

  • Nucleotide sequences only (DNA); protein sequences are not accepted.
  • Sequence names must be unique within the file.
  • At least one sequence must contain 10,000 or more bases other than N.

Repeat masking

Vipsania reads the soft masking of a genome, its lowercase nucleotides, as repeats. Good masking helps, but masking that is too aggressive or too sparse misleads the model. Upload your genome with the masking you trust, or without masking (all uppercase).

File integrity verification

The job status page shows an MD5 checksum of the genome file exactly as it was received by the server. You can verify that your file was transferred without corruption by comparing this value against the MD5 you compute locally:

md5sum your_genome.fa.gz

If the two checksums match, the server processed the exact file you uploaded. This is particularly useful when sharing a job link with a collaborator — they can confirm which genome version was used.

Sample data

A small example genome suitable for a test submission is available here:
example_genome.fa.gz (8 scaffolds from the BRAKER example data set; use with the Streptophyta clade model).

Results

Once your job completes, the status page provides download links for the following files:

  • vipsania.gff3.gz — gene predictions in GFF3 format
  • coding.fasta.gz — predicted coding sequences (CDS) in FASTA format
  • proteins.fasta.gz — predicted protein sequences in FASTA format

All files are gzip-compressed. Decompress them with gunzip before loading them into downstream tools that do not accept compressed input:

gunzip vipsania.gff3.gz
gunzip coding.fasta.gz
gunzip proteins.fasta.gz

Many tools (genome browsers, BLAST, BUSCO, etc.) can read .gz files directly — check the documentation of your tool before decompressing.

Clade model

Vipsania is trained per clade, on raw genome sequences only. Select the most specific clade that contains the organism whose genome you are annotating. Because Vipsania never learns from annotations, finding your species or a close relative among the training species of a model is not an argument against using that model.

Clade Suitable for Model ID Locus F1 on this server
(pretrained)
Locus F1 with finetuning
(not offered here)
Alveolata Ciliates, apicomplexans, dinoflagellates (Plasmodium, Tetrahymena, …) sd2zcj7u 0.582 0.613
Amoebozoa Amoebae and slime moulds (Dictyostelium, Entamoeba, …) ezhpj2qm 0.416 0.437
Arthropoda (other than insects) Crustaceans, arachnids, myriapods; for insects use Insecta ihe1jk30 0.217 0.207
Chlorophyta (green algae) Green algae (Chlamydomonas, …) faeijtmk 0.558 0.579
Cnidaria Corals, sea anemones, jellyfish (Nematostella, Hydra, …) b5vtieo0 0.406 0.453
Discoba Kinetoplastids, euglenids, heteroloboseans (Trypanosoma, Leishmania, Naegleria, …) gcra9d9y 0.696 0.696
Echinodermata Sea urchins, sea stars, sea cucumbers sx9zjl7p 0.440 0.444
Fungi Fungi (Saccharomyces, Aspergillus, …) fh1kg88z 0.645 0.660
Insecta Insects (Drosophila, Apis, …) 58hsuobw 0.474 0.537
Nematoda Roundworms (Caenorhabditis, …) zspca2rb 0.517 0.539
Porifera Sponges (Amphimedon, …) 7bjiexcu 0.416 0.480
Rhodophyta (red algae) Red algae (Porphyra, Cyanidioschyzon, …) 6kmw3wme 0.301 0.405
Spiralia Molluscs, annelids, flatworms v5ej8oyt 0.262 0.374
Stramenopiles Diatoms, brown algae, oomycetes (Phaeodactylum, Phytophthora, …) r6p9z9jw 0.413 0.426
Streptophyta (land plants and relatives) Land plants and charophyte algae (Arabidopsis, rice, …) j9m0cdmk 0.579 0.602
Tunicata Tunicates (Ciona, …) hcehc7ff 0.494 0.523
Vertebrata Fish, amphibians, reptiles, birds, mammals etb1go6q 0.327 0.406
Other eukaryotes (none of the above) Eukaryotes that belong to none of the clades listed above cg6grhms 0.416 0.436

The F1 columns are the average locus-level F1 over the test species of each clade, measured by the Vipsania authors against reference annotations (details). If your organism belongs to none of the clades, choose Other eukaryotes; that model was trained on exactly such species.

Finetuning is not available

Since Vipsania is unsupervised, it can be trained further on the very genome it is about to annotate (--finetune). The Vipsania authors recommend this, and in their benchmark it improved the average locus F1 in almost every clade (compare the two F1 columns above).

This web server does not finetune, and there is no option to enable it. Finetuning is a GPU training run per job, and we do not have the GPU capacity to offer that as a public service. All jobs are annotated with the pretrained clade model. If you need the additional accuracy that finetuning may provide, please run Vipsania on your own hardware; see the tutorial.

E-mail address

Provide a valid e-mail address so that we can send you:

  • A confirmation message when your job is accepted and queued.
  • The results (download links) when the job completes successfully.
  • An error notification if the job fails.

Your e-mail address is deleted as soon as the result or error notification has been sent. If sending fails for any reason, it is deleted at the latest after 45 days (when all job files are removed anyway). It is never used for any purpose other than result notification. See the Data Privacy page for full details.

We do not verify e-mail addresses at submission time; please ensure the address is correct, or you will not receive your results.

No personalized human genomic data

By checking this box you confirm that the genome you are uploading is not a personalized human genome sequence — i.e. it does not contain genomic data derived from an identifiable individual.

Reference assemblies (e.g. GRCh38) and non-human genomes are not affected by this restriction. Uploading personalized human genomic data (e.g. a whole-genome sequencing result from a patient or research participant) is prohibited because this service does not provide the safeguards required for the processing of sensitive personal health data under EU GDPR.

Before submitting a job you must confirm that you have read and agree to our Data Privacy Protection Declaration.

The data processed during your job submission are:

  • Your e-mail address: deleted immediately after the result or error notification is sent; at the latest after 45 days.
  • The genome FASTA file you upload: retained for 45 days after job completion, then automatically deleted.
  • The result files produced by Vipsania: retained for 45 days, then automatically deleted.

No data are shared with third parties. The server is operated by the Bioinformatics Group at Universität Greifswald, Germany, and is subject to EU GDPR regulations.