We published rankings we could not support, and the controls that show it are ones we should have run before selling anything. We withdrew every ranking from sale while we ran the controls below. Sixty-eight have since been restored; twenty-four remain withheld because their signature cannot support a reproducible ranking. Nothing had been delivered to a customer — we checked the fulfilment log and there were no purchases of a repurposing report by anyone outside the company.
The comparison we ran. Twenty-three indications in our corpus hold two independently derived expression signatures. Comparing the twelve top-ranked candidates from each, 14 pairs shared no compound at all, 6 shared one, and 3 shared two. Against the 3,173 pairs built from different indications — which share nothing 60.1% of the time — the same-indication pairs are indistinguishable (60.9% share nothing; Mann-Whitney p = 0.92 two-sided, Fisher p = 1.00). Some pairs of entirely unrelated diseases share six of twelve candidates, which is a measure of how much a top-twelve reflects the compound library rather than the disease.
Why that is not a failed replication. In all 23 pairs one signature was deep (median 931 genes shared with the 978-gene L1000 landmark set) and the other shallow (median 49); two of them rest on mouse data our own species check now blocks. With one side of every comparison that thin, an indistinguishable result was the expected outcome before the test was run. We described this as a failed replication. It is not. It is a test that could not have succeeded, and we should not have sold against it. Whether the method replicates on adequate data is unknown: no indication in this corpus holds two independent deep signatures.
A control we described wrongly, and are withdrawing. We previously said on this page that we shuffle a signature's own gene labels to build “a signature carrying no disease information at all”. That was wrong twice over. The code draws random genes from the whole landmark frame, not from the signature. And when we did run the control we had described — scrambling a signature's values among its own genes, destroying the disease information while keeping its size identical — it passed, as decisively as the real signature did. What our permutation control actually measures is whether a subsample of a signature reproduces that same signature's ranking: stability under subsampling, not validity. At the time we wrote this we had no control that distinguished an informative signature from an uninformative one. We have since built one.
The specificity control, and what it shows. Take a signature and permute its values among its own genes: the gene set is untouched, the depth is untouched, the distribution of values is untouched, and only the assignment of value to gene — the disease content — is destroyed. Then re-rank. Across six signatures, overlap with the real top twelve fell to zero of twelve in every case, a change at least as large as substituting an unrelated disease. So the ranking responds to the disease measurement rather than to which genes happen to be on the assay. Keeping the direction of each gene's change while discarding its magnitude retains roughly a third to a half of the list, so both contribute.
What that does not settle. This is a specificity check and it is the only control we have passed. It does not show that these candidates are biologically correct. It does not show the list would survive an independent signature of the same indication — which remains untestable here, because no indication in this corpus holds two independent deep signatures. And it is not a substitute for laboratory confirmation. Rankings that clear the gate are sold as computational hypotheses on that basis and no stronger.
The gene-count floor, described accurately. Our score is a rank correlation computed only over the genes a disease signature shares with the landmark set. Measured across six signatures, a subsampled signature reproduced its own full ranking in 99.2% of draws at 200 shared genes, 91.7% at 150, and 80.8% at 100. At 50 the result could no longer be separated from the permuted control at all. That is a floor for a stable ranking, not a reliable one. We had been publishing at 50; we now withhold below 300 — noting that 200 is the deepest point we actually measured, so 300 is a conservative margin rather than a measurement. Of 118 reports, 47 are withheld: 41 for gene count, 6 for resting on mouse or rat signatures whose symbols had been mapped to human by uppercasing, which is a crude substitution and not an ortholog mapping.
Three reports withdrawn entirely. Three vertical-level entries (MASH / T2D / obesity, Longevity / senescence, ALK+ NSCLC) had no underlying disease signature at all. A naming mismatch meant they rendered a candidate list that was empty and a gene-overlap figure of “0” that was never measured. They are unpublished rather than repaired, because there is no data behind them to repair.
Candidates are already-approved drugs. For each disease we score a library of compounds against L1000 transcriptional profiles, filter out compounds whose signal is explained by general cytotoxicity rather than a specific effect, and prioritise by clinical maturity using ChEMBL. We then cross-check ClinicalTrials.gov to see whether the drug-indication pair has already been tried.
A measured rather than imputed readout on the genes carrying a candidate's rationale; a cell-type-matched profile where the current one is a proxy; or a registered trial that appears after our last refresh. We would rather revise a ranking than defend one.
Repurposing is the one part of our work with no composition-of-matter to protect — the compounds are already approved. That makes the hypotheses publishable where our novel work is not. What is ownable is the method-of-use position on a new indication: biomarker-selected populations, formulation and 505(b)(2) paths, and combinations. If a candidate here is relevant to a programme you run, that is the conversation to have.