--- license: agpl-3.0 library_name: methscope tags: - dna-methylation - methylation - wgbs - single-cell - imputation - deconvolution - cell-type-annotation --- # MethScope — pretrained models Pretrained models for [**methscope-cli**](https://github.com/zhou-lab/methscope-cli): ultra-fast analysis of sparse DNA methylomes via MRMP (Most Recurrent Methylation Pattern) encoding. Each model is a **single self-contained bundle** that already carries its MRMP feature definition (and labels / signature), so a query `.cg` runs directly — no separate annotation files. ## Fetch & run ```sh # fetch one model (or --local-dir models to grab all) hf download zhou-lab/methscope hg38_celltype.clfx --local-dir . methscope classify query.cg hg38_celltype.clfx # cell type / label methscope deconv mixture.cg hg38_65celltypes.refx # cell-type proportions (NNLS) methscope upscale -o out.cg hg38_wg.updecx query.cg # genome-wide CpG upscaling methscope inspect hg38_sex.clfx # framework, labels, features ``` The file extension names the model's **role** — `.clfx` classifier, `.refx` deconvolution reference, `.updecx` upscale decoder. All three are the same `MSBNDL1` container and are detected by magic, never by name. Classifiers used to be `.ubjx`, which described a payload format (a UBJSON booster) rather than a role and had already stopped being accurate for the non-booster frameworks; `.ubjx` is still accepted everywhere and nothing needs renaming locally. Small query `.cg` test fixtures (typed cells, a deconvolution mixture, an upscale input + truth) live in the GitHub repo [`zhou-lab/methscope_data`](https://github.com/zhou-lab/methscope_data) under `test/`, alongside this catalog and the MRMP-construction reproducibility archive. ## Models | file | task | framework | labels / notes | |------|------|-----------|----------------| | `hg38_wg.updecx` | whole-genome CpG upscaling | `UPDEC2` | all 29,401,795 hg38 CpGs; the primary upscaler (2.8 GB) | | `hg38_10k1.updecx` | CpG upscaling (single block) | MLP decoder | block 10k1 (10,000 CpGs); small demo model | | `hg38_celltype.clfx` | cell-type annotation | `violation` | 33 human cell types; **no fitted parameter** — see below | | `mm10_celltype.ubjx` | cell-type annotation (brain) | xgboost | 41 mouse-brain cell types; rename to `mm10_celltype_brain.clfx` pending a Liu 2021 rebuild | | `hg38_sex.clfx` | sex prediction | logistic | Female, Male (XCI `Xa_hi`/`Xa_lo` markers) | | `hg38_65celltypes.refx` | cell-type deconvolution (NNLS) | refx | 65 types = 58 Zhou + 7 Loyfer organ/blood; split MRMP, 15,300 patterns | `methscope inspect ` prints the exact framework, full label list, and (for linear models) per-feature weights. Classifiers carry a required framework `kind` mark (`xgboost` / `violation` / `threshold` / `logistic`); `classify` rejects an unmarked bundle. Upscale decoders (`.updecx`) run via `upscale`, deconvolution references (`.refx`, `kind=refx`) via `deconv`. ## The `violation` framework `hg38_celltype.clfx` has **no trained parameters at all**. Every MRMP already carries a binstring — the methylation state each reference class takes at that pattern — so the reference *is* a per-class prediction. The model just asks which class the query contradicts least: ``` E1 = patterns the class expects methylated; E0 = expects unmethylated w_j = sqrt(n_cpg_j) v1 = Σ_{j∈E1, observed} w_j·[β_j ≤ 0.5] / Σ_{j∈E1, observed} w_j v0 = Σ_{j∈E0, observed} w_j·[β_j > 0.5] / Σ_{j∈E0, observed} w_j predict argmin_class (v1 + v0), needing ≥20 observed patterns per side ``` `confidence` is the **margin** to the runner-up, not a probability. It is calibrated (accuracy 0.757 / 0.820 / 0.899 at margin ≥ 0.02 / 0.05 / 0.10) and, unlike a softmax, it collapses toward zero when *no* class fits — which is how the model signals that the query's true type is missing from the reference. Records too sparse to clear the minimum are reported `NA` rather than guessed. Held out **across cohorts** (each class's reference and its test samples come from different studies), 113 donor-resolved pseudobulks at ≥100k covered CpGs: **accuracy 0.842, macro-F1 0.853** — against 0.873 / 0.868 for an xgboost model on the identical MRMP. It trades ~3 points of normal-tissue accuracy for behaviour that holds up on tumours. **Single-cell tumours**, where it earns that trade. Bian is split by anatomical site because that is the whole point — a colorectal cancer cell sitting in liver must still read as colon: | cohort | GI fraction | top call | |---|---|---| | Bian 2018 CRC, **liver metastases** | **100.0%** | `Enterocyte.(Colon)` | | Bian 2018 CRC, **lymph metastases** | **100.0%** | `Enterocyte.(Colon)` | | Bian 2018 CRC, primaries (L/R colon, rectum) | 97.1–100.0% | `Enterocyte.(Colon)` | | Bian 2018 adjacent normal (within-patient control) | 100.0% | `Enterocyte.(Colon)` | | Gaiti 2019 CLL | — | `B.Cell` | | Chaligne 2021 glioma | — | `Oligodendrocyte` | Nothing in the reference anticipates the metastatic case, and no bulk cohort can separate origin from site the way these cells do. It also resists the mitotic-age confound that dominates trained models on this feature space: the "most-hypomethylated-class sink" sending globally demethylated tumours to `Acinar.(Pancreas)` drops from 30.0% of TCGA calls to 2.3%, with no residualisation and no tumour data anywhere in its construction. Accuracy is coverage-dependent: 0.163 below 8k covered CpGs, 0.719 at 8–32k, plateauing near 0.85 above ~130k. **Do not use below ~30k covered CpGs.** Limitations: `Macrophage` (0.27) and `Monocyte` (0.74) bleed into `Dendritic.Cell` — the macrophage reference here is BLUEPRINT's, which is monocyte-derived and differentiated *in vitro*, a poor match for tissue macrophages. `T.Cell.CD8` leaks to `CD4` at 27%. Every remaining error is a within-lineage confusion; none cross lineages. The atlas is also incomplete, and a missing origin surfaces as a **low margin** rather than an error — glioma has the lowest median margin of any cohort tested, because neither OPC nor astrocyte is a class. Margins are not comparable between models with different class counts, so recalibrate any threshold per model. For a bulk tumour, deconvolve rather than classify: a classifier assigns one label to a mixture. Rebuild from any MRMP artifact, with no training data, in about 7 seconds: ```sh methscope classify-train --framework violation -p 1000 -o out.clfx ref.mrmp ``` See the [methscope-cli documentation](https://zhou-lab.github.io/methscope-cli/) for the full command reference and runnable examples.