Skip to content

Examples

Scenario-based examples for the most common RolyPoly workflows/commands. Note - command help messages are updated on a more frequent basis, so for full option lists, run rolypoly <command> --help.


filter-reads and assemble

Background

tl;dr preset choice

How your RNA was prepared for sequencing is usually the main thing that determines which preset to use. Three factors tend to matter most: 1. whether your library was poly-A selected, 2. whether you want to filter potential host cDNA/mRNA and other contaminants, and 3. whether your sample came from enriched/purified virions ("virome"), where you can often skip rRNA removal and sometimes use lighter quality filtering.

  • If your library is poly-A selected: RolyPoly includes a poly-A trimming step in filter-reads.
  • If you want host filtering: provide a host reference genome to filter-reads and assemble (RolyPoly will try to mask likely viral regions in that reference before using it for filtering).
  • If your sample is virome-enriched: you can often skip rRNA and host removal, use lighter quality filtering, and sometimes enable deduplication steps (depending on your goals).

Note: "host" removal can also remove EVEs or integrated viral sequences, so it may not be ideal if you are interested in non-RNA viruses or endogenous viral elements related to your study.


From sample to reads: the wet-lab pipeline

Before we get "reads", the biological material goes through a series of wet-lab steps that strongly shape what the sequencing data look like. Understanding that chain helps with picking in-silico settings and avoiding unexpected read loss.

The steps and methods below are usually chosen based on study goals and budget. Many introduce bias, but also reduce cost or increase throughput. For example, rRNA depletion kits do not equally target all domains of life, but they can substantially reduce the sequencing depth needed.

In most non-targeted genomics or metagenomics studies, the starting material can be anything from a single host species and its infecting virus, to a complex mixture spanning multiple domains of life (table 1). The "meta" prefix (e.g. metagenomics, metatranscriptomics) generally implies a mixed community as the source.

Regardless of source type, material is usually processed through a mix of chemical, enzymatic, and physical steps depending on study focus. The tables below summarize common steps across the wet-lab pipeline: sample handling before nucleic acid extraction (table 2), extraction itself (table 3), post-extraction processing (table 4), and library preparation (table 5). For RNA virus-centric/transcriptomic work, this usually means RNA extraction, often with rRNA depletion or poly-A selection, followed by reverse transcription to cDNA. Direct RNA sequencing is also possible (e.g. Oxford Nanopore), though still less common in many workflows.

Note: this overview is intentionally practical and not exhaustive.

Table 1: Common sample sources for RNA virus discovery

Source What it contains Notes for virus discovery
Monoculture / host cells / tissue Predominantly host RNA; viral RNA if infected High host rRNA background expected. Consider including a host reference for filtering if target virus(es) can be distinguished from host sequences - note that some viruses have endogenized relatives (EVEs) or an integrated stage (e.g. retroviruses).
Environmental (water / soil / sediment / biofilm / atmospheric dust...) Mixed community (bacteria, archaea, eukaryotes, viruses...) No single host reference; rRNA dominates.
Environmental concentrate (e.g. filter retentate 0.2–3 µm, ultracentrifugation) Similar to raw environmental sources, but pre-processed to enrich or deplete size fractions Can be used for community-level expression analysis (metatranscriptomics) as well as virus discovery. Depending on pre-processing, free virions may be depleted, meaning virus discovery may reflect intracellular viral RNA rather than extracellular virions.
VLPs / protected environmental nucleic acids (e.g. ultracentrifugation / sucrose gradient / FeCl₃ precipitation / 0.1 µm filtrate / TFF) Enriched sub-cellular fraction; virions, membrane vesicles, mobile genetic elements; some host debris DNase/RNase treatment prior to extraction can greatly reduce host background, but low yields may necessitate amplification (e.g. RCA, MDA), which can introduce chimeras and skew abundances.

Once the sample is collected, it usually goes through pre-extraction processing. These steps act on physical or chemical properties of the sample (size, density, protection status, membrane/protein integrity) and can strongly affect what ends up in the extracted material. For our use case, one key distinction is whether the sample was further enriched for viruses (e.g. VLP-focused workflows, often called "viromics"). Some steps also serve multiple practical purposes, such as buffer exchange, inhibitor removal, or transport/storage stabilization.

A practical note on filtration: the steps listed in table 2 can be used in combination, and it is entirely possible to use the retentate from one filter as the input to a second filtration step, or to collect the filtrate instead - the choice depends on which size fraction is of interest.

Table 2: Common pre-nucleic acid extraction processing steps

Processing Step What it does Notes for virus discovery
Centrifugation (low speed / clarification) Removes large debris and intact cells by pelleting Standard first step in many protocols; supernatant is carried forward
Homogenization / bead beating Disrupts tissue or cell-rich material into a homogenate Releases intracellular contents including host RNA; increases host background
Filtration (large pore, e.g. 5–20 µm) Removes large debris, most eukaryotic cells, and multicellular material Helps reduce host cell material; virions and small prokaryotes pass through freely
Filtration (small pore, e.g. 0.2–0.45 µm) Retains bacteria and larger particles; virions pass into filtrate Can deplete prokaryotic cells; filtrate is enriched for virions and small VLPs
Tangential flow filtration (TFF) Concentrates particles within a size range Useful for large volume samples (e.g. seawater); retains virions while removing small molecules
Ultracentrifugation / sucrose gradient Separates particles by size and density Can enrich for VLP/virion fractions; may co-pellet host debris and membrane vesicles
Chemical treatment (e.g. FeCl₃ precipitation) Precipitates virions or nucleic acid-associated particles Low-cost concentration method for large volumes; efficiency varies by virus type
Proteinase K treatment Degrades proteins in the sample Can improve nucleic acid yield and purity by removing protein-nucleic acid complexes; may also disrupt non-enveloped virion capsids if not controlled carefully
Chloroform / organic solvent treatment Denatures and removes lipids and proteins Commonly used in phenol-chloroform extraction; also used pre-extraction to disrupt enveloped viruses and remove host membrane material - note this will also disrupt enveloped virions
DNase / RNase treatment (pre-extraction) Degrades unprotected extracellular nucleic acids in the sample Reduces free host nucleic acid contamination; nucleic acids inside virions or vesicles are protected and retained

Many of these steps are combined within a single protocol or commercial kit. Some steps also serve multiple roles - for example, a buffer exchange step may simultaneously remove inhibitors while concentrating the sample.


Nucleic acid extraction is the step that liberates nucleic acids from biological material. The method you choose can introduce bias in yield, purity, and which nucleic acid species are recovered. For RNA virus discovery, RNA extraction is usually the default, though total nucleic acid (TNA) extraction is also used when both DNA and RNA are of interest (for example, mixed RNA/DNA virus surveys or integrated proviral sequence detection).

Table 3: Common nucleic acid extraction methods

Method What it recovers Notes for virus discovery
TRIzol / phenol-chloroform extraction Total RNA (and DNA if desired) Robust and widely used; recovers a broad size range of RNA; requires careful handling of hazardous reagents; co-extracted DNA can be removed by DNase treatment if needed
Silica column-based RNA extraction (e.g. Qiagen RNeasy, Zymo) Total RNA above a size threshold (typically >200 nt) Convenient and fast; small RNAs (e.g. siRNAs, miRNAs) may be lost unless a specific small RNA kit is used; yields can be low for dilute or complex samples
Total nucleic acid (TNA) extraction Both DNA and RNA Useful when both RNA and DNA viruses are of interest; downstream DNase or RNase treatment can separate the pools if needed
Small RNA extraction / enrichment RNA below ~200 nt (miRNA, siRNA, piRNA) Specifically targets small RNA species; relevant for studies of RNA silencing or small RNA-based virus detection; full-length viral genomes will be depleted
Direct lysis (e.g. in low-input or single-cell protocols) Variable; depends on lysis conditions Minimises handling loss; may co-extract inhibitors; often paired with whole-transcriptome amplification

After extraction, a second round of processing is often used to reshape the nucleic acid pool by removing unwanted species or adjusting yield/purity before library prep. These steps act on the molecules directly (not the physical sample) and can have a large effect on what gets sequenced.

Table 4: Common post-nucleic acid extraction processing steps

Processing Step What it does Notes for virus discovery
DNase treatment (on-column or in-solution) Degrades residual DNA from the extracted nucleic acid pool Important for RNA-focused studies to reduce DNA carryover; critical if reverse transcription is used, to avoid amplifying DNA templates
Ribosomal RNA (rRNA) depletion Removes rRNA using probe-based hybridization and depletion (e.g. RiboZero, NEBNext) rRNA is often the dominant fraction in total RNA (commonly around 80-90%+); depletion increases the proportion of informative transcripts, but probe sets may not cover non-model organisms well
Poly-A selection Captures polyadenylated RNA using oligo-dT beads Enriches for eukaryotic mRNA and some RNA viruses that have poly-A tails; depletes rRNA, bacterial RNA, and non-polyadenylated viral RNA
Size selection Retains RNA fragments within a target size range Can be used to enrich for small RNAs such as siRNAs or miRNAs; may deplete full-length viral genomes
RNA concentration / cleanup Removes salts, enzymes, or other inhibitors (e.g. column cleanup, ethanol precipitation) Important for downstream enzymatic steps; low-yield samples risk loss during cleanup
Amplification (e.g. RCA, MDA) Increases the total amount of nucleic acid Necessary for low-yield samples; can introduce chimeric reads, amplification bias, and skewed abundances - complicates quantification

Library preparation converts processed nucleic acid into sequencing-ready molecules. For RNA libraries, this begins with reverse transcription to cDNA; after that, the workflow broadly resembles DNA library prep. Choices at this stage, especially fragmentation, priming strategy, and strand specificity, can strongly affect how viral sequences are represented and interpreted downstream.

When working with multiple samples (for example, time points, locations, or tissue types), it helps to plan sample-level barcoding/indexing before library prep starts. Each sample can get a unique index (or dual-index combination), so libraries can be pooled in one sequencing run and demultiplexed later. This is standard and usually lowers per-sample cost. But prep decisions still matter: if samples are pooled too early (for example at the RNA stage), sample identity is lost. Keeping a clear index-to-sample map is essential for meaningful downstream comparisons. In practice, indexing strategy is often constrained by the sequencing center to avoid collisions with other projects in the same run. Many users therefore receive already-demultiplexed data; if not, you may also see index FASTQs in addition to R1/R2 files.

Table 5: Common library preparation steps

Step What it does Notes for virus discovery
Reverse transcription (RT) Converts RNA to cDNA using reverse transcriptase Required for RNA sequencing on most short-read and some long-read platforms; primer choice (random hexamers, oligo-dT, or gene-specific) affects coverage and representation. In non-targeted approaches, sequence-specific primers shouldn't be used
Second-strand synthesis Converts single-stranded cDNA to double-stranded cDNA (dsDNA) Enables ligation-based library preparation; strand information may be lost depending on the method used
RNA/cDNA fragmentation Breaks nucleic acid into shorter fragments suitable for sequencing Fragmentation method (chemical, enzymatic, heat, or sonication) can affect coverage uniformity and introduce biases; may be performed on RNA prior to RT or on cDNA afterward
End repair & A-tailing Blunts fragment ends and adds an adenosine overhang Prepares fragments for adapter ligation; standard step in most Illumina library prep workflows
Adapter ligation Ligates platform-specific adapters to fragment ends Adapters contain primer binding sites and optional barcodes/indices; ligation efficiency affects library complexity
Indexing / barcoding Adds unique index sequences to each sample's library Enables multiplexing of multiple samples in a single sequencing run; dual indexing reduces index-hopping artifacts. See note above on sample tracking strategy
PCR amplification Amplifies the adapter-ligated library Excess PCR cycles reduce library complexity and introduce duplicates - particularly problematic for quantification of viral abundance
Size selection (post-ligation) Selects fragments within a target insert size range (e.g. 150–500 bp) Removes adapter dimers and very short/long fragments; affects read length distribution and insert size
Strand-specific / directional library prep Preserves information about which strand the RNA was transcribed from Important for correctly orienting viral RNA segments, identifying antisense transcription, and reducing ambiguity in de novo assemblies

From sample to FASTQ: the processing chain

flowchart TD
    A[Sample\ncells / retentate / VLP pellet] --> B[RNA extraction\ne.g. TRIzol, column-based]
    B --> C{Enrichment or depletion?}
    C -->|Total RNA + ribo-depletion\nRibo-Zero, RNase H-based| D[ribo-depleted RNA\n★ recommended for broad virus discovery]
    C -->|Poly-A selection\noligodT capture| E[poly-A-enriched mRNA]
    C -->|No treatment\ntotal RNA| F[total RNA\nhigh rRNA content]
    D & E & F --> G[RNA fragmentation\nenzymatic or physical shearing\ntarget ~150–400 bp]
    G --> H[Reverse transcription\nrandom hexamers or oligo-dT]
    H --> I[End repair + dA-tailing]
    I --> J[Adapter ligation\nP5 + P7 Illumina adapters]
    J --> K[PCR amplification\n6–15 cycles]
    K --> L[Size selection\nSPRI beads or gel\nsets insert size distribution]
    L --> M[Paired-end sequencing\ntypically 2×75 bp or 2×150 bp]
    M --> N[Raw FASTQ\nR1 + R2 per sample]

### From digitized data to in-silico processing

Important: the goal of in-silico processing is not to remove as many reads as possible. The goal is to reduce technical noise and increase biological signal for your specific question.

This part is inherently goal-dependent. Trimming, filtering, masking, host subtraction, normalization, and deduplication are tools, not goals by themselves. If a step does not clearly improve signal for the biological target you care about, it is often better to keep it conservative or skip it.

Pipeline, tool, and parameter choices should follow the study target, but they could also be informed by everything that happened before digitization (sample source, enrichment strategy, extraction chemistry, library prep, expected contaminants, expected insert size, and sequencing platform).

Examples:

  • If you also care about DNA viruses, avoid aggressive host/DNA removal.
  • If you know exactly which adapter sequence was used, start with that specific adapter set instead of broad default databases (some of which may match real viral sequences).
  • If your focus is differential expression (host), preserve host reads and avoid filtering choices that distort host transcript abundance.
  • If your focus is low-level variation (for example SNPs or quasispecies/subspecies structure), avoid error correction, over-normalization, and aggressive deduplication that can flatten minor variants.
  • If your focus is recovering major viruses and improving genome completeness, stronger cleanup and downsampling reads with over represented kmers may actually make assembly easier.

It is also completely fine to run multiple in-silico "branches" from the same sample, each with different pipeline choices. For example, you can use the original raw reads (or lightly processed reads) for quantification, while using a normalized subset for assembly-oriented recovery.

Many recent workflows also combine strategies, such as pooling selected samples, running multiple assemblers and merging/curating outputs, applying k-mer abundance-based read normalization, and in some cases adding haplotyping/strain-resolution analyses after the fact.

Common in-silico steps mapped to filter-reads

The list below follows a common read-processing order and maps each stage to the internal filter-reads step names. Each step includes examples for how to skip it (--skip-steps), when it is skipped by presets, and/or how to tune behavior with --override-parameters.

Decision flow for read-processing branches

flowchart TD
    A[Raw FASTQ] --> B{Need host or known DNA removal?}
    B -->|yes| C[filter_known_dna]
    B -->|no| D[skip filter_known_dna]

    C --> E{Need rRNA depletion in silico?}
    D --> E
    E -->|yes| F[decontaminate_rrna]
    E -->|no| G[skip decontaminate_rrna]

    F --> H{Use identified-DNA filtering?}
    G --> H
    H -->|yes| I[filter_identified_dna]
    H -->|no| J[skip filter_identified_dna]

    I --> K[dedupe]
    J --> K

    K --> L[trim_adapters]
    L --> M{Poly-A library?}
    M -->|yes| N[trim_polya_tails]
    M -->|no| O[skip trim_polya_tails]

    N --> P[remove_synthetic_artifacts]
    O --> P
    P --> Q[entropy_filter]

    Q --> R{Short inserts / overlap expected?}
    R -->|yes| S[error_correct_1]
    S --> T[error_correct_2]
    T --> U[merge_reads]
    R -->|no| V[skip overlap-heavy steps]

    U --> W[quality_trim_unmerged]
    V --> W
    W --> X[final dedupe + outputs]

    X --> Y{Single branch or multiple branches?}
    Y -->|single| Z[One downstream path]
    Y -->|multiple| AA[Branch e.g. quantification vs assembly]

1) Known DNA filtering: filter_known_dna

Used for host or known DNA contaminant subtraction when you provide -D/--known-dna. If --known-dna is not provided, this step is automatically skipped.

# Use a custom known DNA reference
rolypoly filter-reads -i reads/ -o filtered/ \
  -D host_or_contaminant.fasta

# Skip known DNA filtering explicitly
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps filter_known_dna

# Tune matching strictness
rolypoly filter-reads -i reads/ -o filtered/ -D host.fasta \
  --override-parameters '{"filter_known_dna": {"k": 31, "mincovfraction": 0.8, "hdist": 0}}'

2) rRNA decontamination: decontaminate_rrna

Uses packaged rRNA references (SILVA + NCBI masked sets).

# Default run (with rRNA filtering)
rolypoly filter-reads -i reads/ -o filtered/

# Skip rRNA filtering
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps decontaminate_rrna

# Also skipped by this preset
rolypoly filter-reads -i reads/ -o filtered/ \
  --preset all_virus_metag

# Make rRNA filtering stricter/looser
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"decontaminate_rrna": {"mincovfraction": 0.7, "k": 31}}'

3) Identified DNA filtering: filter_identified_dna

This step uses the rRNA stats profile to fetch candidate genomes and filter likely host/DNA reads.

# Run with identified-DNA filtering
rolypoly filter-reads -i reads/ -o filtered/

# Skip identified-DNA filtering
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps filter_identified_dna

# Already skipped by these presets
rolypoly filter-reads -i reads/ -o filtered/ --preset fast
rolypoly filter-reads -i reads/ -o filtered/ --preset all_virus_metat
rolypoly filter-reads -i reads/ -o filtered/ --preset all_virus_metag

# Tune filtering sensitivity
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"filter_identified_dna": {"mincovfraction": 0.8, "k": 31}}'

4) Deduplication (early pass): dedupe

dedupe appears here in the main processing chain, and another dedupe pass is run at final output stage.

# Skip the early dedupe stage in the main chain
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps dedupe

# Tune dedupe aggressiveness
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"dedupe": {"passes": 1, "s": 0}}'

# strict preset increases dedupe aggressiveness
rolypoly filter-reads -i reads/ -o filtered/ --preset strict

5) Adapter trimming: trim_adapters

Adapter trimming runs after early decontamination and before quality trimming.

# Skip adapter trimming (usually not recommended)
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps trim_adapters

# Tune adapter trim behavior
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"trim_adapters": {"k": 23, "mink": 11, "hdist": 1, "minlen": 20}}'

# If you want to use only a known custom adapter set, pre-trim externally,
# then skip internal adapter trimming in rolypoly.
cutadapt -a file:my_adapters.fa -A file:my_adapters.fa \
  -o pretrim_R1.fq.gz -p pretrim_R2.fq.gz reads_R1.fq.gz reads_R2.fq.gz

rolypoly filter-reads -i pretrim_R1.fq.gz,pretrim_R2.fq.gz -o filtered/ \
  --skip-steps trim_adapters

6) Poly-A tail trimming: trim_polya_tails

Useful for poly-A selected libraries; disabled by default unless preset/flag enables it.

# Enable poly-A tail trimming explicitly
rolypoly filter-reads -i reads/ -o filtered/ --trim-polya

# Also enabled by this preset
rolypoly filter-reads -i reads/ -o filtered/ --preset poly_a_selected

# Skip poly-A trimming
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps trim_polya_tails

# Tune poly-A trimming
rolypoly filter-reads -i reads/ -o filtered/ --trim-polya \
  --override-parameters '{"trim_polya_tails": {"trimpolya": 18, "minlen": 20}}'

7) Synthetic artifact filtering: remove_synthetic_artifacts

Targets synthetic/control artifacts.

# Skip synthetic artifact filtering
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps remove_synthetic_artifacts

# Tune k-mer matching
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"remove_synthetic_artifacts": {"k": 31}}'

8) Low-complexity filtering: entropy_filter

Filters very low-complexity reads (mostly homopolymers / low-entropy sequence).

# Skip entropy filtering
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps entropy_filter

# Tune entropy thresholds
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"entropy_filter": {"entropy": 0.01, "entropywindow": 30}}'

9) Overlap-based correction stage 1: error_correct_1

Most useful when paired reads overlap (short inserts). On single-end input, this step is auto-skipped.

# Skip stage 1 correction
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps error_correct_1

# fast preset already skips error_correct_1
rolypoly filter-reads -i reads/ -o filtered/ --preset fast

# Tune stage 1 behavior
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"error_correct_1": {"mix": "t", "ordered": "t"}}'

10) Correction stage 2: error_correct_2

Second correction stage. This one is not auto-skipped for single-end by default.

# Skip stage 2 correction
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps error_correct_2

# fast preset already skips error_correct_2
rolypoly filter-reads -i reads/ -o filtered/ --preset fast

# Tune stage 2 behavior
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"error_correct_2": {"passes": 1, "reorder": true}}'

11) Overlap merge: merge_reads

Merges overlapping pairs. Auto-skipped for single-end input.

# Skip merge stage
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps merge_reads

# Tune merge behavior
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"merge_reads": {"mix": "f"}}'

12) Quality trimming of unmerged reads: quality_trim_unmerged

Late-stage quality trimming after correction/merge steps.

# Skip quality trimming
rolypoly filter-reads -i reads/ -o filtered/ \
  --skip-steps quality_trim_unmerged

# Tune trim stringency
rolypoly filter-reads -i reads/ -o filtered/ \
  --override-parameters '{"quality_trim_unmerged": {"trimq": 12, "minlen": 25}}'

Final output stage note

After the main chain above, filter-reads runs a final dedupe pass on merged/interleaved outputs.

Key decisions that affect your RolyPoly preset

  • ribo-depletion vs. poly-A selection - a major choice that can affect RNA virus recovery (see section below for evidence and caveats).
  • Insert size - set by the size-selection step. Typical targets are 150–400 bp for Illumina. Short inserts (< 2×read length) produce overlapping read pairs; very short inserts bleed into adapter sequence.
  • Amplification method - standard PCR introduces PCR duplicates. SISPA or random-primed amplification (used in some VLP protocols) introduces additional chimeric read artifacts (Kugelman et al., 2017).

Library preparation and viral recovery

Many RNA viruses, including phages and most negative-sense, ambisense, and segmented RNA viruses, are not polyadenylated, so enrichment strategy can strongly affect what gets recovered. In a direct comparison, poly(A)-selected libraries yielded viral reads but were insufficient for complete recovery of a non-polyadenylated virus genome, while ribo-depleted total RNA performed substantially better (Visser et al., 2016; PMID: 27250973). In practice, total RNA plus ribo-depletion is often a good starting point for discovery-oriented workflows.

Ribo-depletion also has limits. Probe performance depends on sequence match, so depletion can be less effective in non-model or mixed communities (Kim et al., 2019; PMID: 31783730). Computational rRNA filtering is therefore commonly used as a second layer (Zhou et al., 2018; PMID: 29444661), but very aggressive filtering can discard informative reads. Dovrolis et al. showed that rRNA-focused sorting can capture rRNA-virus chimeras and reduce viral support for assembly (PMID: 34370725).

Independent RNA-seq evidence also supports treating host-virus chimeras cautiously: host-virus chimeric events can be infrequent and largely artifactual, consistent with RT template switching during library construction (Yan et al., 2021; PMID: 33980601). For preprocessing, this argues for conservative thresholds and careful interpretation of putative chimeric reads instead of blanket removal based on single-feature matches.

Subjective note: I have also observed some rRNA-virus chimeric contigs in published virus discovery works and general metatranscriptomes (see Table S6, sheet "rRNA_summary" in Neri et al., 2022), and the full extent of host-virus chimera is probably underappreciated. I am not sure I can recommend keeping chimeric reads though - I assume these are not biological and often arise from template switching after nucleic acid extraction - and properly trimming only the host portion of chimeric reads is not trivial (IMO).

RolyPoly preprocessing currently focuses on Illumina short-read libraries (poly-A selected, ribo- depleted, or unselected total RNA). Long-read/direct-RNA support is a separate ongoing area.


Sequencing artifacts and their computational signatures

Each step between nucleic acid extraction and the sequencer can introduce artifacts.

Paired-end reads and insert size

The insert is the actual cDNA fragment between the two Illumina adapters. Both ends are sequenced, producing R1 (reads the forward strand) and R2 (reads the reverse complement strand).

Adapter-P5 ─── INSERT (cDNA fragment) ─── Adapter-P7
    │                                           │
    └──► R1 reads →                 ← R2 reads ◄┘

Typical layout (insert longer than read length, no overlap):

5'─[P5]─[R1 ████████████████]─ · · · gap · · · ─[████████████████ R2]─[P7]─3'
         read 1 (fwd)           unsequenced middle           read 2 (rev-comp)
         └────────────────────────────────────────────────────────────────────┘
                              insert size (e.g. 300 bp)

Insert size is set by the size-selection step in library prep. Typical Illumina runs use 2×75 bp or 2×150 bp reads; inserts are usually 150–400 bp. The ratio of insert to read length determines which artifacts appear.

Short inserts, overlapping reads, and adapter contamination

As insert size decreases relative to read length, two problems arise:

Case 1 - Overlapping reads (insert slightly shorter than 2× read length):

5'─[P5]─[R1 ████████████]─[████ overlap ████]─[████████████ R2]─[P7]─3'

R1: ████████████████████                        (reads left → right)
R2:             ████████████████████            (reads right → left, RC)
              ←── overlap region ──►

BBMerge can merge overlapping pairs into a single, longer, more accurate read.
This is common for short viral genomes where short inserts are preferred.
Case 2 - Adapter contamination (insert shorter than read length):

5'─[P5]─[fragment]─[P7]─3'   (very short insert)

R1: [fragment sequence ████][P7-adapter sequence ░░░░░░░░░░]
                              ↑ adapter "bleed-in"
R2: [fragment RC ████][P5-adapter RC ░░░░░░░░░░]

Without adapter trimming, these tails can look like spurious sequence and reduce assembly quality.
Adapter trimmers (e.g., BBDuk, cutadapt, Trimmomatic) typically remove these by matching known
adapter sequence k-mers, and most also let you provide custom adapter sequences.

PCR duplicates vs optical duplicates

Both appear as multiple reads with identical sequence, but their causes differ:

PCR duplicates                    Optical duplicates
──────────────────────────────    ──────────────────────────────
Same molecule amplified           Single cluster on the flowcell
multiple times during             that spreads or is misread as
library prep PCR:                 two adjacent clusters:

Flowcell surface:                 Flowcell surface:

 ● cluster A  (seq: ACGTACGT)      ●●  cluster A+B touching
 ● cluster B  (seq: ACGTACGT)          (seq: ACGTACGT for both)
 ● cluster C  (seq: ACGTACGT)
 (scattered anywhere on tile)     ← physically adjacent, same tile

Cause: amplification bias         Cause: optical/diffusion artifact
Worse with: more PCR cycles,      Worse with: patterned flowcells
  low-input / VLP libraries         (NovaSeq, NextSeq 2000)
Removed by: sequence identity     Removed by: sequence identity
  (any duplicate-removal tool)      AND cluster proximity

Both can inflate abundance estimates and mislead assembly coverage. In metatranscriptomes, though, highly expressed transcripts are expected, so aggressive duplicate removal can also discard real biology. In RolyPoly, deduplication is done as exact-sequence removal (seqkit rmdup) rather than coordinate-based deduplication.

Virus–host chimeric reads

Template switching during reverse transcription, or ligation of fragmented RNA molecules during library prep, can join viral and host RNA into a single read:

Origin molecules:

  ─────[viral genomic RNA segment]─────3'
                                        ↕ RT jumps (template switching)
  5'─────[host mRNA ─────────────]─────

Resulting chimeric cDNA / read:

  5'─[viral sequence ███████████][host sequence ░░░░░░░░░░░]─3'
      maps to viral genome ─┘         maps to host ────┘

Consequence 1: chimeric read is filtered by host-removal step
               → viral portion is lost along with the host portion

Consequence 2: chimeric read enters assembly
               → chimeric contig spanning virus + host
               → false annotations or missed virus

Consequence 3: rRNA–virus chimeras (described in Dovrolis et al., 2021; PMID: 34370725)
               → rRNA-sorting tools bin the chimera as rRNA
               → virus reads lost; assembly of complete genome fails

This is one reason RolyPoly keeps a permissive rRNA filter by default: strict rRNA filtering can discard the viral component of mixed/chimeric reads (Dovrolis et al., 2021; PMID: 34370725). The same caveat applies to host subtraction, where very strict mapping thresholds may remove reads that only partially match host sequence.

Practical implications for preset choice

  • If broad RNA virus recovery is the priority, a practical starting point is total RNA with ribo-depletion and moderate rRNA filtering.
  • If libraries are multiplexed on patterned-flowcell instruments, non-redundant dual indexing can help reduce index-swap cross-talk (Costello et al., 2018; PMID: 29739332).
  • If inserts are short and read overlap is high, read merging often improves effective read quality and assembly input (Bushnell et al., 2017; PMID: 29073143).
  • If low-input amplification was used, it is usually safer to handle duplicate/chimera filtering conservatively and validate candidates by remapping support (Kugelman et al., 2017; PMID: 28182717).

Background references (PubMed-checked)

  • Visser M, Bester R, Burger JT, Maree HJ. Next-generation sequencing for virus detection: covering all the bases. Virol J. 2016;13:85. doi:10.1186/s12985-016-0539-x. PMID: 27250973.
  • Dovrolis N, Kassela K, Konstantinidis K, et al. ZWA: Viral genome assembly and characterization hindrances from virus-host chimeric reads; a refining approach. PLoS Comput Biol. 2021;17(8):e1009304. doi:10.1371/journal.pcbi.1009304. PMID: 34370725.
  • Yan B, Chakravorty S, Mirabelli C, et al. Host-Virus Chimeric Events in SARS-CoV-2-Infected Cells Are Infrequent and Artifactual. J Virol. 2021;95(15):e00294-21. doi:10.1128/JVI.00294-21. PMID: 33980601.
  • Neri U, Wolf YI, Roux S, et al. Expansion of the global RNA virome reveals diverse clades of bacteriophages. Cell. 2022;185(21):4023-4037.e18. doi:10.1016/j.cell.2022.08.023. PMID: 36174579.
  • Kim IV, Ross EJ, Dietrich S, et al. Efficient depletion of ribosomal RNA for RNA sequencing in planarians. BMC Genomics. 2019;20(1):909. doi:10.1186/s12864-019-6292-y. PMID: 31783730.
  • Zhou Q, Su X, Jing G, Chen S, Ning K. RNA-QC-chain: comprehensive and fast quality control for RNA-Seq data. BMC Genomics. 2018;19(1):144. doi:10.1186/s12864-018-4503-6. PMID: 29444661.
  • Costello M, Fleharty M, Abreu J, et al. Characterization and remediation of sample index swaps by non-redundant dual indexing on massively parallel sequencing platforms. BMC Genomics. 2018;19(1):332. doi:10.1186/s12864-018-4703-0. PMID: 29739332.
  • Bushnell B, Rood J, Singer E. BBMerge - Accurate paired shotgun read merging via overlap. PLoS One. 2017;12(10):e0185056. doi:10.1371/journal.pone.0185056. PMID: 29073143.
  • Kugelman JR, Wiley MR, Nagle ER, et al. Error baseline rates of five sample preparation methods used to characterize RNA virus populations. PLoS One. 2017;12(2):e0171333. doi:10.1371/journal.pone.0171333. PMID: 28182717.
  • Adiconis X, Borges-Rivera D, Satija R, et al. Comparative analysis of RNA sequencing methods for degraded or low-input samples. Nat Methods. 2013;10(7):623-629. doi:10.1038/nmeth.2483. PMID: 23685885.
  • He S, Wurtzel O, Singh K, et al. Validation of two ribosomal RNA removal methods for microbial metatranscriptomics. Nat Methods. 2010;7(10):807-812. doi:10.1038/nmeth.1507. PMID: 20852648.

Preset quick-reference

roll preset Library preparation Filter preset Assembly preset
rna_virus (default) RNA virus metatranscriptome: rRNA removal, host + identified-DNA filter rna_virus_metat rna_virus
ribodepleted Total RNA ribo-depleted: stricter rRNA removal (mincovfraction=0.7) total_rna_ribodepleted rna_virus
poly_a Poly-A selected mRNA: polyA tail trim, stricter quality trim poly_a_selected metatranscriptome
all_virus_metat All-virus metatranscriptome / RNA virome: relaxed rRNA filter, skips identified-DNA filter all_virus_metat rna_virus
DNA_virus DNA virome / metagenomics: skips rRNA and identified-DNA filtering all_virus_metag metag (metaSPAdes only)
complete Any — maximum sensitivity; runs all three assembler modes rna_virus_metat complete
fast Any — quick preview; skips error correction and identified-DNA filter fast fast

When unsure, start with --preset rna_virus. Check read-count retention in the log before committing to a full run on many samples.


End-to-end pipeline (roll)

The roll command runs the full discovery pipeline in one call: read filtering → assembly → contig filtering → marker search → nucleotide search → annotation. Pick the --preset that matches your library preparation.

Viral RNA metatranscriptome (default)

Ribo-depleted total RNA from an environmental sample; expected to contain RNA viruses. This is the default preset, so --preset rna_virus can be omitted.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_out/ \
  --threads 16 --memory 64g \
  --preset rna_virus

With host/contaminant removal

Provide a FASTA of the host genome (or any expected DNA contaminant) with -D. The assembly step will also filter contigs that match the host.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_out/ \
  --threads 16 --memory 64g \
  --preset rna_virus \
  --host host_genome.fasta

Total RNA, ribo-depleted library

Use ribodepleted for stricter rRNA removal (mincovfraction=0.7).

rolypoly roll \
  --input reads/ \
  --output-dir rp_out/ \
  --threads 16 --memory 64g \
  --preset ribodepleted

Poly-A selected mRNA library

Enables polyA tail trimming and uses rnaSPAdes+MEGAHIT assembly.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_out/ \
  --threads 16 --memory 64g \
  --preset poly_a

DNA virome / metagenomics (no rRNA or identified-DNA filtering)

Suitable for DNA-based viromes or metagenomic libraries where you do not want rRNA or identified-DNA filtering applied.
NOTE: this isn't really rolypoly forte, this preset is just for convenience if you want to run a general pipeline that WON'T remove DNA data, or harm assembly of potential hosts.

rolypoly roll \
  --input reads/ \
  --output-dir rp_out/ \
  --threads 16 --memory 64g \
  --preset DNA_virus

Maximum sensitivity

Runs all three assembler modes (metaSPAdes + rnaviralSPAdes + MEGAHIT). Slowest, but highest chance of recovering divergent or low-abundance viruses.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_out/ \
  --threads 32 --memory 128g \
  --preset complete

Quick preview with --mini

Subsamples the input before running the pipeline; forces the fast assembly preset. Useful for a rapid sanity-check before committing to a full run.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_preview/ \
  --threads 8 --memory 16g \
  --preset rna_virus \
  --mini

Override individual sub-presets

Use --filter-preset and/or --assembly-preset to mix-and-match independently of --preset.

# Strict read filtering, but fast assembly
rolypoly roll \
  --input reads/ \
  --output-dir rp_out/ \
  --filter-preset strict \
  --assembly-preset fast

Resume a partial run

WARNING ! THIS IS NOT FULLY TESTED YET. Use at your own risk. --skip-existing skips any step whose output directory already exists.

rolypoly roll \
  --input reads_R1.fq.gz,reads_R2.fq.gz \
  --output-dir rp_out/ \
  --preset rna_virus \
  --skip-existing

Read filtering (filter-reads)

Paired-end reads, default settings

rolypoly filter-reads \
  -i reads_R1.fq.gz,reads_R2.fq.gz \
  -o filtered_reads/ \
  -t 8 -M 16g

Directory of FASTQ files

All FASTQ files in the directory are processed; paired files are matched by base name.

rolypoly filter-reads \
  -i raw_reads/ \
  -o filtered_reads/ \
  -t 8

With a preset

# RNA virus metatranscriptome (lenient quality trim, rRNA removal at mincovfraction=0.6)
rolypoly filter-reads -i reads/ -o filtered/ --preset rna_virus_metat

# Total RNA ribo-depleted (stricter rRNA removal mincovfraction=0.7)
rolypoly filter-reads -i reads/ -o filtered/ --preset total_rna_ribodepleted

# Poly-A selected library (enables polyA trimming, stricter quality trim)
rolypoly filter-reads -i reads/ -o filtered/ --preset poly_a_selected

# All-virus metatranscriptome (relaxed rRNA filter, skips identified-DNA filter)
rolypoly filter-reads -i reads/ -o filtered/ --preset all_virus_metat

# All-virus metagenomics (skip rRNA + identified-DNA filters entirely)
rolypoly filter-reads -i reads/ -o filtered/ --preset all_virus_metag

# Fast (skip error correction and identified-DNA filter)
rolypoly filter-reads -i reads/ -o filtered/ --preset fast

With host removal

rolypoly filter-reads \
  -i reads_R1.fq.gz,reads_R2.fq.gz \
  -o filtered_reads/ \
  -D host_genome.fasta \
  --preset rna_virus_metat

Override a specific step parameter

# Raise rRNA coverage threshold and use a stricter quality trim
rolypoly filter-reads \
  -i reads/ -o filtered/ \
  --preset rna_virus_metat \
  --override-parameters '{"decontaminate_rrna": {"mincovfraction": 0.8}, "quality_trim_unmerged": {"trimq": 15}}'

Skip specific steps

# Run everything except error correction
rolypoly filter-reads \
  -i reads/ -o filtered/ \
  --skip-steps error_correct_1 --skip-steps error_correct_2

Assembly (assemble)

From a directory of filtered reads, default settings

rolypoly assemble \
  -id filtered_reads/ \
  -o assembly_out/ \
  -t 16 -M 64g

With a preset

# RNA virus: rnaviralSPAdes + MEGAHIT, broad k-mer range, rmdup post-processing
rolypoly assemble -id filtered_reads/ -o assembly_out/ --preset rna_virus

# Metatranscriptome: rnaSPAdes + MEGAHIT
rolypoly assemble -id filtered_reads/ -o assembly_out/ --preset metatranscriptome

# Metagenomics: metaSPAdes (for DNA-based libraries)
rolypoly assemble -id filtered_reads/ -o assembly_out/ --preset metag

# Fast: MEGAHIT only, narrow k-mer range
rolypoly assemble -id filtered_reads/ -o assembly_out/ --preset fast

# Complete: all three assembler modes
rolypoly assemble -id filtered_reads/ -o assembly_out/ --preset complete

Explicit library specification

# Paired-end
rolypoly assemble \
  --paired-end 1 reads_R1.fq.gz reads_R2.fq.gz \
  -o assembly_out/ -t 8 -M 32g

# Multiple libraries mixed
rolypoly assemble \
  --paired-end 1 lib1_R1.fq.gz lib1_R2.fq.gz \
  --merged 2 lib2_merged.fq.gz \
  -o assembly_out/ -t 8 -M 32g

Run only rnaviralSPAdes

rolypoly assemble \
  -id filtered_reads/ \
  -o assembly_out/ \
  -A spades_rnaviral

Override k-mer settings

rolypoly assemble \
  -id filtered_reads/ -o assembly_out/ \
  --preset rna_virus \
  --override-parameters '{"megahit": {"k-min": 27, "k-max": 99, "k-step": 12}}'

Skip post-processing deduplication

rolypoly assemble \
  -id filtered_reads/ -o assembly_out/ \
  --preset rna_virus \
  --skip-steps post_processing

Modular step-by-step workflow

The roll command chains these steps automatically, but each can be run independently for finer control or to slot into an existing pipeline.

# 1. Filter reads
rolypoly filter-reads \
  -i raw_reads/ -o filtered_reads/ \
  -t 16 -M 32g --preset rna_virus_metat

# 2. Assemble
rolypoly assemble \
  -id filtered_reads/ -o assembly/ \
  -t 16 -M 64g --preset rna_virus

# 3. Filter assembled contigs against a host reference
rolypoly filter-contigs \
  -i assembly/final_assembly.fasta \
  --host host_genome.fasta \
  -o assemblies/filtered_assembly.fasta \
  -t 8

# 4. Search for viral marker genes (RdRps, genomad)
rolypoly marker-search \
  -i assemblies/filtered_assembly.fasta \
  -o marker_results/ \
  -t 16

# 5. Nucleotide-level search against known RNA virus databases
rolypoly virus-mapping \
  -i assemblies/filtered_assembly.fasta \
  -o virus_hits.tab \
  -t 16

# 6. Annotate candidate contigs
rolypoly annotate \
  -i assemblies/filtered_assembly.fasta \
  -o annotation/ \
  -t 16

# Search all supported databases (RdRp HMMs + geNomad)
rolypoly marker-search \
  -i contigs.fasta -o marker_out/ -t 8

# Search only the RdRp database
rolypoly marker-search \
  -i contigs.fasta -o marker_out/ --database rdrp -t 8

Virus nucleotide mapping (virus-mapping)

rolypoly virus-mapping \
  -i contigs.fasta \
  -o virus_hits.tab \
  -t 8

Shrink / subsample reads

# Random subsample to 50 000 reads (for a quick test)
rolypoly shrink-reads \
  -i reads_R1.fq.gz,reads_R2.fq.gz \
  -o sampled/ \
  --subset-type random --sample-size 50000

# Coverage-normalise with bbnorm (better for paired data)
rolypoly shrink-reads \
  -i reads_R1.fq.gz,reads_R2.fq.gz \
  -o sampled/ \
  --subset-type bbnorm