MMTAX¶
Summary¶
Assign ICTV taxonomy to virus contigs from protein matches made by MMseqs2 or DIAMOND.
Both search backends feed the same post-search classifier:
- Keep matches within
--toppercent of each protein's best bitscore. - Collapse duplicate hits assigned to the same taxid to their best score.
- Walk down the ICTV ranks using weighted votes, first for each protein and then across the proteins on each contig.
Each selected rank must be a child of the previously selected rank, so the
result is always one lineage that exists in the taxonomy. A match labelled only
at a shallow rank contributes there but is neutral in deeper-rank votes. Thus a
large weight assigned only to root cannot erase lower-weight family evidence,
while conflicting family-resolved matches still compete normally. This rule is
generic and does not privilege any particular virus lineage.
The vote can be weighted by bitscore (default) or negative log E-value. This makes backend comparisons less confounded by different built-in ORF and LCA implementations.
Usage¶
With the built-in database:
Nucleotide contigs are passed through the existing pyrodigal-rv predictor. To
reuse proteins from an earlier RolyPoly step, supply the FASTA and an explicit
two-column protein<TAB>contig map:
rolypoly mmtax \
--input contigs.fasta \
--proteins predicted_orfs.faa \
--protein-map protein_to_contig.tsv \
--output taxonomy.tsv
--infer-protein-map is available for RolyPoly/pyrodigal-style headers ending
in an ORF number, but an explicit map is safer for arbitrary protein FASTA
files.
For direct protein input, use --query-type protein. Without a map, each
protein is classified as its own output query. Supplying --protein-map instead
groups those proteins back onto contigs before weighted assignment.
For a custom database, --taxdump is mandatory:
rolypoly mmtax \
--input contigs.fasta \
--database /path/to/ictv_nr_db/ictv_nr_db \
--taxdump /path/to/ictv_taxdump \
--backend mmseqs \
--output taxonomy.tsv
A custom MMseqs2 database needs the _mapping and _taxonomy sidecars from
mmseqs createtaxdb. A custom DIAMOND database must be built with --taxonmap
and --taxonnodes. A plain FASTA file is not a taxonomy database.
Shared search controls¶
--sensitivity accepts a number from 1 through 8 or the corresponding names:
faster, fast, mid-sensitive, normal, sensitive, more-sensitive,
very-sensitive, and ultra-sensitive. normal is the default and maps to
MMseqs2 sensitivity 4 or DIAMOND's unmodified default preset.
--min-aln-len is passed directly to MMseqs2. For DIAMOND it is applied to the
tabular hits after search and before taxonomy assignment. RolyPoly applies
--top once, as the same post-search bitscore-window filter for both backends,
before calculating the rank-aware assignment.
--tax-lineage 1 writes names, --tax-lineage 2 writes taxids (the default),
and --tax-lineage 0 omits the compact lineage. Named ICTV-rank columns are
always retained for reporting.
Output¶
The headered TSV contains the assigned query, taxid, rank, taxon_name,
overall support, informative_fraction, counts of assigned proteins and
retained hits, lineage, backend, and method. Each ICTV rank also has a name
and conditional support column, for example family and family_support.
support describes agreement among matches that resolve to the assigned rank;
informative_fraction reports how much compatible retained-hit weight actually
resolved that deeply. A high-support call with a low informative fraction should
therefore be treated as tentative.
The best individual retained alignment is reported separately as
best_match_target, its taxid, name and rank, and its bitscore, E-value,
identity, and alignment length. This is evidence, not necessarily the final
assignment: for example, the best-scoring hit may be labelled only at root
while other rank-resolved hits support a family.
Breadth columns make that distinction more auditable:
proteins_assigned / total_proteinsandprotein_hit_fractionreport how many predicted genes on the contig had retained matches.aligned_residues / total_protein_residuesandresidue_alignment_fractionreport the union of covered query-amino-acid positions, so overlapping database matches are not counted repeatedly.projected_aligned_nt / genome_lengthandprojected_alignment_genome_fractionreport the union of those amino-acid intervals projected through single-CDS GFF coordinates. They remain empty for protein-only mappings or proteins whose genomic coordinates cannot be projected safely.
RolyPoly's HTML report discovers mmtax.tsv in an output directory and adds a
taxonomy table plus a family-composition chart. The roll command runs taxonomy
by default; use --skip-steps taxonomy to omit it.
rvmt and nvpc are reserved database names. They will remain disabled until
their profiles have been enriched with compatible ICTV taxids.
Caveats¶
Defaults can produce overly specific assignments¶
Current defaults in the command implementation are:
| Control | Default | Meaning |
|---|---|---|
--backend |
mmseqs |
Protein similarity search backend |
--sensitivity |
normal |
MMseqs2 sensitivity 4; DIAMOND default preset |
--evalue |
1e-5 |
Maximum hit E-value |
--identity |
0.1 |
Minimum aligned amino-acid identity: 10% |
--min-aln-len |
30 |
Minimum alignment length in amino acids |
--top |
10 |
Retain hits within 10% of the best protein bitscore |
--weight |
bitscore |
Weight for taxonomic voting |
--majority |
0.5 |
Minimum weighted support used in rank-aware voting |
These permissive search thresholds and voting rules can assign a divergent or
fragmentary contig too deeply. Support measures agreement among retained,
rank-informative reference matches; it is not a calibrated probability of
membership. A high support value, especially with a low informative_fraction,
does not establish that an assignment satisfies the biology of that taxon.
Similarity assignment is not taxonomic demarcation¶
Conceptually, mmtax belongs to the similarity-search-to-taxonomy family of methods often described as BLAST-to-LCA. Its current implementation is more specifically a lineage-consistent, rank-aware weighted vote at protein and contig levels, rather than a strict LCA of every retained hit.
It does not apply actual rank- and taxon-specific demarcation criteria, nor calibrated approximations to them. There is no universal RdRp amino-acid identity cutoff for family, genus or species membership. An aligned-region identity is also not interchangeable with identity across a complete RdRp, polyprotein or genome. The relevant comparison and criteria vary across virus groups and ranks; see the ICTV explanation of taxonomic classification.
For example, a contig may receive a family label because its best represented matches all belong to that family, even when its RdRp similarity falls below the range or criterion appropriate for membership. Conserved polymerase homology can support a broader evolutionary relationship without establishing membership of that family. Genome organization and accessory-protein content may provide additional conflicting evidence. Their absence from a partial contig, however, is not proof that the complete virus lacks those genes.
Treat assignments as hypotheses to review using the relevant ICTV report, alignment extent, phylogenetic placement, genome completeness and gene content. Raising generic identity or voting thresholds alone does not implement the appropriate demarcation criteria. See also the scientific background and report-chart caveats.
Known bugs¶
The limitations above concern the classifier's interpretation and calibration; they are not a claim that it implements formal taxon demarcation incorrectly. Visual/search integration issues are tracked under report known bugs.