Data availability
Download
Data files used by this site: the sample-level metadata of the reference cohort, the social niche embedding matrix, the trait probabilities and the reference cohort scores.
Files
Loading the file list…
File descriptions
disease_sample_metadata.tsv — Sample-level metadata for the reference cohort of the dysbiosis score: 10,276 human faecal 16S rRNA samples (4,386 controls and 5,890 cases) from 13 diseases. Tab-separated, one row per sample; an empty cell is a value the original study did not report.
| Column | Contents |
|---|---|
| sample | Sequencing run accession |
| study | BioProject accession |
| title | Title of the source study |
| diagnosis | Diagnosis as reported by the study |
| group | 0 for controls, 1 for cases |
| disease_name, disease_name_ab | Disease name and its abbreviation |
| age, sex, bmi, host_body_mass_index, country, subject_id | Host information |
| region, pcr_primers, platform | Amplified 16S region, primers and sequencing platform |
| site | Sampling site |
| Abbreviation | Disease | Abbreviation | Disease |
|---|---|---|---|
| AS | Ankylosing spondylitis | IBS | Irritable bowel syndrome |
| ASD | Autism spectrum disorder | MS | Multiple sclerosis |
| BD | Bipolar disorder | OB | Obesity |
| CAD | Coronary artery disease | PD | Parkinson's disease |
| CRC | Colorectal cancer | SZ | Schizophrenia |
| GD | Graves' disease | T2DM | Type 2 diabetes mellitus |
| IBD | Inflammatory bowel disease |
sne_vectors.txt.gz — The 14,093 × 100 embedding matrix in GloVe text format, one OTU per line. OTU identifiers are SILVA 138.2 accession numbers with the alignment span, as used in the OTU tables.
traits_proba.f16.bin — Random forest probabilities for all classes of every trait, for each labelled OTU, in the layout described in traits.json. These values are used to draw the distribution plots in the atlas.
traits.json — Trait metadata: the class order of each trait, its offset into traits_proba.f16.bin, and its leave-one-phylum-out AUC.
ref_scores.json — Sorted model scores of the 10,276 reference samples, split into controls and cases. Percentiles are computed from this distribution.