Interactive atlas
Gut Microbial Social Niche Atlas
Each point represents one OTU, positioned according to its ecological embedding. Two OTUs lie close together when they share co-occurrence partners, that is when the same taxa surround them in a community, rather than when they occur together; that shared company is what makes their ecological niches similar, independent of phylogenetic relatedness. Select an OTU to compare its ecological and phylogenetic neighbours side by side.
Examples: (genus) · (species) · (OTU) · , (300-bp V4 reads)
Names follow the SILVA 138.2 taxonomy throughout this site, and matching is on family, genus and species. An OTU identifier is that OTU's SILVA reference accession followed by the aligned region, accession.start.stop, as used in the OTU tables. Only 1,245 of the 14,093 OTUs (9%) carry a species name in SILVA, so a genus is usually the more reliable query.
Loading atlas data…
Methods and interpretation
Layout. Point positions are a UMAP projection of the 100-dimensional ecological embeddings. Distances in two dimensions are approximate and should not be interpreted quantitatively; the ecological neighbour list gives cosine similarities in the full 100-dimensional space.
Names and identifiers. Every name on this site comes from the SILVA 138.2 SSU Ref NR99 taxonomy, so the ranks shown for an OTU are that one system rather than a mixture of sources. SILVA 138.2 uses the current phylum names, which differ from the older ones still common in the literature: Bacillota (formerly Firmicutes), Bacteroidota (Bacteroidetes), Pseudomonadota (Proteobacteria) and Actinomycetota (Actinobacteria). An OTU identifier is the accession of the OTU's SILVA reference sequence followed by the region of it that is aligned, in the form accession.start.stop; the identifier HM007585.1.1335, for example, is bases 1 to 1,335 of GenBank accession HM007585. These are the identifiers the OTU tables use, and the ones the dysbiosis score expects. A species name is given only where the SILVA record names an isolate: that is 1,245 of the 14,093 OTUs. The others were sequenced from uncultured material and carry no species in SILVA, so they are reached by genus, by identifier, or by searching their sequence.
The two neighbour lists. Ecological neighbours are ranked by the cosine similarity of the social niche embeddings, so 1.000 would be an identical co-occurrence pattern across the approximately 210,000 samples. Phylogenetic neighbours are ranked by patristic distance on the SILVA 138.2 reference tree — the sum of the branch lengths along the path between the two OTUs, in substitutions per site — so 0 would be the same position in the tree. The two numbers answer different questions and are not on one scale; compare values within a column, not across the pair.
Trait prediction. Traits are predicted by random forest classifiers trained on SNE vectors, with Traitar genome annotations as labels. Performance is estimated by leave-one-phylum-out cross-validation, in which all members of the test phylum are excluded from training. This is stricter than random splitting, and the resulting AUC is shown next to every prediction.
Sequence search. Pasted or uploaded sequences are aligned on the server against the representative sequences of all 14,093 OTUs at 97% identity, and are deleted after the search. A short amplicon may be equally similar to several OTUs; in that case, all matching OTUs are listed.
Hidden traits. Predictions for traits with a cross-validated AUC below 0.65 are hidden by default because their predictive performance is low; they can be shown at the bottom of each OTU panel. BacDive measurements and Traitar genome annotations are always shown.