Data availability

Download

Data files used by this site: the sample-level metadata of the reference cohort, the social niche embedding matrix, the trait probabilities and the reference cohort scores.

Files

Loading the file list…

File descriptions

disease_sample_metadata.tsv — Sample-level metadata for the reference cohort of the dysbiosis score: 10,276 human faecal 16S rRNA samples (4,386 controls and 5,890 cases) from 13 diseases. Tab-separated, one row per sample; an empty cell is a value the original study did not report.

ColumnContents
sampleSequencing run accession
studyBioProject accession
titleTitle of the source study
diagnosisDiagnosis as reported by the study
group0 for controls, 1 for cases
disease_name, disease_name_abDisease name and its abbreviation
age, sex, bmi, host_body_mass_index, country, subject_idHost information
region, pcr_primers, platformAmplified 16S region, primers and sequencing platform
siteSampling site
AbbreviationDiseaseAbbreviationDisease
ASAnkylosing spondylitisIBSIrritable bowel syndrome
ASDAutism spectrum disorderMSMultiple sclerosis
BDBipolar disorderOBObesity
CADCoronary artery diseasePDParkinson's disease
CRCColorectal cancerSZSchizophrenia
GDGraves' diseaseT2DMType 2 diabetes mellitus
IBDInflammatory bowel disease

sne_vectors.txt.gz — The 14,093 × 100 embedding matrix in GloVe text format, one OTU per line. OTU identifiers are SILVA 138.2 accession numbers with the alignment span, as used in the OTU tables.

traits_proba.f16.bin — Random forest probabilities for all classes of every trait, for each labelled OTU, in the layout described in traits.json. These values are used to draw the distribution plots in the atlas.

traits.json — Trait metadata: the class order of each trait, its offset into traits_proba.f16.bin, and its leave-one-phylum-out AUC.

ref_scores.json — Sorted model scores of the 10,276 reference samples, split into controls and cases. Percentiles are computed from this distribution.