Tutorial

This tutorial walks you through preparing your input data and interpreting the gene prediction output produced by Vipsania.

No finetuning on this server. Vipsania can be finetuned on the genome it is about to annotate (vipsania annotate --finetune), and its authors recommend doing so. This web server does not offer finetuning, because we do not have the GPU capacity for it: every job is annotated with the pretrained clade model as it is. Finetuning may make the annotation somewhat more accurate. If you need that, please run Vipsania locally with --finetune; a GPU is strongly recommended for finetuning.

1. Prepare your genome

Vipsania works with any eukaryotic genomic FASTA file. For best results, follow these guidelines:

  • File size limit. Gzip-compressed files (genome.fasta.gz) up to 500 MB are accepted, and the genome may contain at most 300 Mbp.
  • Sequence IDs must be unique within the file. Avoid spaces or special characters in FASTA headers; unique sequence IDs up to the first whitespace are required for correct matching with the GFF3 output.

2. Choose a clade model

Select the most specific clade that contains your organism:

Dropdown labelRecommended for
AlveolataCiliates, apicomplexans, dinoflagellates (Plasmodium, Tetrahymena, …)
AmoebozoaAmoebae and slime moulds (Dictyostelium, Entamoeba, …)
Arthropoda (other than insects)Crustaceans, arachnids, myriapods; for insects use Insecta
Chlorophyta (green algae)Green algae (Chlamydomonas, …)
CnidariaCorals, sea anemones, jellyfish (Nematostella, Hydra, …)
DiscobaKinetoplastids, euglenids, heteroloboseans (Trypanosoma, Leishmania, Naegleria, …)
EchinodermataSea urchins, sea stars, sea cucumbers
FungiFungi (Saccharomyces, Aspergillus, …)
InsectaInsects (Drosophila, Apis, …)
NematodaRoundworms (Caenorhabditis, …)
PoriferaSponges (Amphimedon, …)
Rhodophyta (red algae)Red algae (Porphyra, Cyanidioschyzon, …)
SpiraliaMolluscs, annelids, flatworms
StramenopilesDiatoms, brown algae, oomycetes (Phaeodactylum, Phytophthora, …)
Streptophyta (land plants and relatives)Land plants and charophyte algae (Arabidopsis, rice, …)
TunicataTunicates (Ciona, …)
VertebrataFish, amphibians, reptiles, birds, mammals
Other eukaryotes (none of the above)Eukaryotes that belong to none of the clades listed above

If your organism is not covered by any clade, choose Other eukaryotes. The model is applied as it is; this server does not finetune it on your genome (see Local installation and finetuning).

3. Submit the job

  1. Go to Submit Job.
  2. Upload your genome file (or use the browse button to locate it).
  3. Select the clade model.
  4. Optionally enter your email address to receive a notification when the job finishes, and tick the consent checkbox.
  5. Click Submit Job.

You will be redirected to a status page. Bookmark it; it updates automatically every 60 seconds. You will also receive an email when the job finishes or fails.

The status page also displays the MD5 checksum of your uploaded genome file (computed from the raw compressed bytes). Compare it with the output of md5sum your_genome.fa.gz to confirm that the file was uploaded intact.

4. Runtime expectations

After submission, your job enters a wait queue before execution begins. Wait time depends on current server load and can range from a few minutes to several hours. The status page shows whether your job is queued or running. Depending on resource availability, jobs may run on CPU or GPU hardware; GPU execution is substantially faster, so actual runtimes can vary significantly from the estimates below.

HardwareGenome sizeMeasured runtime
GPU (one A100)27–73 Mbp2–7 min
CPU (32 cores)27–73 Mbp1–7 h

Runtime grows with genome size. Because any job may run on CPU, this server accepts genomes of up to 300 Mbp. Jobs time out after 72 hours.

5. Interpret the output

Vipsania produces a gene annotation file in GFF3 format. Each predicted gene has one transcript model (mRNA) with annotated exon and CDS features. Coding-sequence and protein-sequence FASTA files are also included. The annotation was produced by the pretrained clade model without finetuning; keep this in mind when you compare it to published Vipsania accuracy figures, which are usually reported with finetuning.

Decompress the result files

All result files are gzip-compressed (.gz). Decompress them with gunzip before passing them to tools that do not accept compressed input:

gunzip vipsania.gff3.gz gunzip coding.fasta.gz gunzip proteins.fasta.gz

Many tools (genome browsers, BLAST, BUSCO, etc.) can read .gz files directly without decompression — check your tool's documentation first.

Example GFF3 output (first few lines):

##gff-version 3 NC_007200.1 Vipsania gene 1925 6467 . + . ID=g1;Name=g1 NC_007200.1 Vipsania mRNA 1925 6467 . + . ID=g1_t1;Name=g1_t1;Parent=g1 NC_007200.1 Vipsania CDS 1925 5948 . + 0 ID=g1_t1.CDS.1;Parent=g1_t1 NC_007200.1 Vipsania exon 1925 5948 . + 0 ID=g1_t1.exon.1;Parent=g1_t1 NC_007200.1 Vipsania CDS 6031 6467 . + 2 ID=g1_t1.CDS.2;Parent=g1_t1 NC_007200.1 Vipsania exon 6031 6467 . + 2 ID=g1_t1.exon.2;Parent=g1_t1

You can load this file directly into genome browsers such as JBrowse or IGV.

6. Local installation and finetuning

This web server never finetunes. We do not have the GPU capacity to run a training job for every submission, so there is no option for it. Finetuning adapts the model to the codon usage, repeat landscape and intron statistics of your species and may make the annotation somewhat more accurate. To get it, or for large genomes, repeated runs, or private data, run Vipsania locally (a GPU is strongly recommended):

python -m pip install vipsania vipsania annotate Fungi genome.fa -o annotation.gff3 --finetune

or with the container image:

singularity pull vipsania.sif docker://gaiusaugustus/vipsania:latest singularity exec --nv -B /path/to/data:/data vipsania.sif \ vipsania annotate Fungi /data/genome.fa -o /data/annotation.gff3 --finetune

Without --finetune, the same commands reproduce what this server does. See the Vipsania GitHub repository for full installation and usage documentation.