Source-linked AI summary
Deriving chemosensitivity from cell lines: Forensic bioinformatics and reproducible research in high-throughput biology
Keith A. Baggerly, Kevin R. Coombes
TL;DR
The paper addresses whether poorly documented microarray-based drug-sensitivity signatures can be reproduced reliably when they are used to predict patient response. It uses forensic reconstruction across related studies and finds simple errors hidden by inadequate documentation, while reporting that correctly applied predictions were no better than chance in the authors’ tests. The authors therefore discuss complete scripts and detailed reporting as safeguards.
Problem
High-throughput studies often lack enough processing documentation for exact reproduction, and concealed errors can affect treatment decisions based on drug-sensitivity signatures.
Method
The authors reconstruct analyses across related papers by examining raw data, reported results, cell-line identities, gene lists, and prediction procedures in several case studies.
Results
Forensic reconstruction identified simple errors across the case studies, while predictions made from NCI60 cell lines without those errors were no better than chance.
Takeaways & Limitations
Simple switches, offsets, and label mix-ups can be hidden by incomplete documentation, so the authors argue that complete scripts will eventually be required alongside raw data.
Takeaways & Limitations
Reconstructing the analyses required combining information across multiple sources, and some conclusions could become unfalsifiable when methods were not detailed.
Abstract
from arXiv · showhide
High-throughput biological assays such as microarrays let us ask very detailed questions about how diseases operate, and promise to let us personalize therapy. Data processing, however, is often not described well enough to allow for exact reproduction of the results, leading to exercises in "forensic bioinformatics" where aspects of raw data and reported results are used to infer what methods must have been employed. Unfortunately, poor documentation can shift from an inconvenience to an active danger when it obscures not just methods but errors. In this report we examine several related papers purporting to use microarray-based signatures of drug sensitivity derived from cell lines to predict patient response. Patients in clinical trials are currently being allocated to treatment arms on the basis of these results. However, we show in five case studies that the results incorporate several simple errors that may be putting patients at risk. One theme that emerges is that the most common errors are simple (e.g., row or column offsets); conversely, it is our experience that the most simple errors are common. We then discuss steps we are taking to avoid such errors in our own investigations.
1. Background.
Microarrays promise personalized treatment, but inadequate documentation can make high-throughput results irreproducible and conceal errors with clinical consequences. The report examines drug-sensitivity signatures, reconstructs how related analyses were produced, and identifies simple errors while discussing safeguards for reproducible research.
- Motivation: Microarrays offer detailed biological measurements and may help identify which patients will respond to standard therapy.The paper illustrates this potential using ovarian cancer, where about 70% of patients respond to standard front-line treatment.
- Reproducibility and risk: Poorly documented data processing can prevent exact reproduction and obscure errors that affect clinical treatment decisions.The authors describe this as especially concerning when signatures are used to allocate patients to clinical-trial treatment arms.
- Signature-based prediction: The examined approach derives drug-sensitivity signatures by selecting sensitive and resistant cell lines, identifying differentially expressed genes, and building a response-classification model.The model is then used to predict patient response from array profiles.
- Subsequent progress: The approach attracted substantial attention after reports of successful predictions for multiple chemotherapeutic agents and combination therapies.It was cited in 212 papers by August 2009 and became the basis for further validation studies and clinical-trial applications.
- Initial claims: An independent reanalysis of doxorubicin data found nearly inverted sensitive/resistant proportions, suggesting that the labels might have been reversed.The reported cohort had 23 sensitive and 99 resistant patients, whereas the underlying datasets contained 94 sensitive and 28 resistant patients.
- Cases examined: The authors examine four cases in detail, selecting examples involving doxorubicin, cisplatin, pemetrexed, combination therapy, and temozolomide.Their choices reflect available prediction information, biologically plausible genes, current treatment guidance, multi-drug treatment, and recency.
2. Case study 1: Doxorubicin.
The doxorubicin case study found that publicly posted data contained reversed training labels, duplicated and inconsistently labeled test samples, and broader classification discrepancies that undermined reproducibility.
- Data acquisition: The posted Adria ALL.txt file contained 22 training columns and 122 test columns, with 99 labeled NR and 23 labeled Resp.These counts matched the numbers reported by Potti et al. (2006).
- Training data sensitive/resistant labels are reversed: Brute-force row correlations matched all 8,958 rows to transformed NCI60 data and identified the 10 resistant and 12 sensitive training cell lines.The recovered ordering reproduced the original heatmap, but the labels were reversed relative to the later website listing.
- Heatmaps show sample duplication in the test data [Figure 1(a)]: Only 84 of 122 test samples were distinct: 60 appeared once, 14 twice, 6 three times, and 4 four times.Some duplicated samples received inconsistent sensitive/resistant labels.
- Heatmaps show sample duplication in the test data [Figure 1(a)]: Columns 32, 66, 89, and 117 were identical, yet were labeled Resp, NR, NR, and NR, respectively.The heatmap showed no clear responder/nonresponder separation and highlighted tied blocks with inconsistent labels.
- At least 3/8 of the test data is incorrectly labeled resistant (Table 3): The n95.doc file contained 95 rows but only 80 distinct samples; 15 duplicates included 6 samples labeled both RES and SEN.Comparing classifications with Holleman et al. (2004) showed that 29–35 sensitive samples and 10 intermediate samples were classified as resistant.
- Communication with the journal elicited a second correction: Poor documentation concealed label reversal and duplicate or mislabeled samples, and these problems survived two explicit corrections.The authors provide code and documentation for the case study in Supplementary File 1.
3. Case study 2: Cisplatin and pemetrexed.
The cisplatin signature failed to separate sensitive and resistant cell lines as reported, but a one-row offset reproduced the published heatmap and exposed indexing, labeling, and platform errors. The pemetrexed signature was likewise exactly matched only after offsetting.
- 3.1. A heatmap using the cisplatin genes shows no separation of the cell lines: The named cisplatin genes produced no clear split between sensitive and resistant cell lines across the 30-line panel.The reconstructed heatmap showed little structure.
- 3.2. A heatmap using offset cisplatin genes shows clear separation of the cell lines: A single-row offset produced a clear separation between sensitive and resistant cell lines.Quantifications from row 98 were used instead of row 97 for an example probeset.
- 3.3. Clustering correlations suggests the cell lines involved: Clustering the offset-gene correlations identified two groups, with 10 lines in one group and 7 in the other, matching the reported subset structure.The resistant group was inferred from the Györffy et al. labels.
- 3.4. Applying binreg perfectly reproduces the reported heatmap: Applying binreg to the inferred cell lines produced an exact match to Hsu et al.’s cisplatin heatmap and identified the genes involved.The matching reconstruction confirmed the cell lines used.
- 3.5. The software produces 41/45 offset genes; the others are the ones explicitly mentioned: The reconstruction matched 41 of 45 cisplatin probesets after offsetting, while four reported outliers included probesets absent from the U133A measurements.The absent probesets were on the U133B platform used neither by the source dataset nor the reconstructed heatmaps.
- 3.6. Pemetrexed: After offsetting, all 85 pemetrexed genes and its published heatmap were matched, identifying 8 resistant and 10 sensitive cell lines.The reconstructed sensitive and resistant groups were explicitly enumerated in the analysis.
- 3.7. Summary: Poor documentation obscured an off-by-one error affecting all reported genes, inclusion of genes from other arrays, and reversal of sensitive/resistant labels.These errors affected the reported signatures examined in the case study.
4. Case study 3: Combination therapy.
The combination-therapy analysis found treatment perfectly confounded with array run date and scanner, while undocumented combination rules produced irreproducible validation results. The cyclophosphamide prediction could not be independently matched and its cell-line sensitivity data showed no differential activity.
- 4.1. Array data contain three run blocks: Three array blocks corresponded to run-date blocks, with treatment perfectly confounded with run date and scanner.The third block contained all TET arrays, which were run on a different scanner from the first two blocks.
- 4.2. Three different combination rules were used: The combination rules were inferred rather than explicitly documented, and all three rules differed and were nonstandard.The authors could not determine which rule had been validated.
- 4.3. We can’t match the accuracy for the best drug/treatment combination: The reported cyclophosphamide ROC curve had AUC 0.943, whereas the reconstructed curve had AUC 0.348.The reported curve exactly matched Bonnefoi et al.’s curve, but the reconstructed curve was qualitatively different.
- 4.4. Sensitivity to cyclophosphamide doesn’t separate the cell lines used: Cyclophosphamide sensitivity did not differentially separate the selected cell lines because cyclophosphamide is a prodrug with no direct effect on cell lines.The selected cyclophosphamide lines matched those used for the pemetrexed signature.
- 4.5. Summary: The study did not mention design confounding, and poor documentation left the scoring computation and cyclophosphamide cell-line selection unclear.The analysis reports these as unresolved aspects of the combination-therapy validation.
5. Case study 4: Temozolomide.
The initially reported temozolomide heatmap did not correspond to the named drug or claimed cell-line panel and was identical to the cisplatin heatmap. A correction replaced it with a substantially different heatmap and gene count.
- 5.1. Individual genes: The initial temozolomide gene list contained 8 probesets higher in resistant lines and 37 higher in sensitive lines, but three genes were listed as higher in both groups.Those three genes were interrogated by multiple probesets.
- 5.2. The initial heatmap matches cisplatin: The reported temozolomide heatmap was identical to the cisplatin heatmap and therefore did not derive from the NCI-60 temozolomide cell lines.The cisplatin heatmap was independently regenerated from the Györffy et al. data.
- 5.3. Journal communication led to a new heatmap with different problems: The correction replaced the initial 45-gene heatmap of 9 resistant and 6 sensitive lines with a 150-gene heatmap of 5 resistant and 5 sensitive lines.The fraction of probesets higher in resistant lines changed from 8/45 to about 110/150.
- 5.4. Summary: Poor documentation left the report combining a heatmap for one drug with a gene list for another, while the results were supported only by visual inspection and counting.The analysis states that these results were not documented further.
6. Case study 5: Surveying cell lines used.
The authors assembled cell-line sensitivity information for ten drugs from twelve sources and reconstructed how sensitive/resistant labels and orientations were assigned. They found label reversals across repeated sources, inconsistencies in cell-line sets, and several drug-specific discrepancies.
- Data sources: Twelve information sources were examined for sensitivity signatures covering ten drugs, including docetaxel, doxorubicin, cisplatin, pemetrexed, and temozolomide.The sources included heatmaps, gene lists, website quantifications, and cell-line lists.
- Inference procedure: Heatmaps and gene lists were matched to binreg outputs to identify contrasted cell-line groups, while label direction was inferred from statements in the relevant papers.Some sources left the direction or exact cell-line identities imprecise.
- Cross-source inconsistencies: Every drug checked more than once showed at least one reversal of sensitive/resistant labeling across information sources.The figure summarizes these changes as color flips across rows.
- Cross-source inconsistencies: Cell-line sets differed for most drugs; cyclophosphamide and pemetrexed were the exception, but their reported gene lists could not be reproduced from the stated cell-line subsets.The cyclophosphamide-producing set was a superset of the pemetrexed set, while the reported cyclophosphamide cell lines were a subset.
- Drug-specific findings: The cisplatin signature used 30 cell lines, while the temozolomide heatmap matched the cisplatin heatmap.The cisplatin cell lines were assembled by Györffy et al. (2006).
- Drug-specific findings: Assuming the August 2008 cell-line orientations were correct, Salter heatmaps were correct for topotecan and doxorubicin but incorrect for fluorouracil and cyclophosphamide.The authors judged Potti heatmaps reversed for topotecan, doxorubicin, and fluorouracil, and could not reproduce the cyclophosphamide heatmap.
7. Discussion.
The discussion argues that forensic reconstruction exposes simple errors hidden by incomplete documentation and that such problems are especially consequential when results guide clinical trials. It recommends complete computational disclosure and documents procedural changes adopted by the authors.
- Scope of findings: The case studies are illustrative rather than exhaustive, and supplementary reports describe additional problems of similar kinds.The authors present the cases as examples of a broader reproducibility problem.
- Common errors: Simple errors can involve annotation mixups, column shifts, confounding, or omitted files, yet incomplete documentation often hides their simplicity.Earlier examples included a one-cell deletion followed by inappropriate shifting of values in one column.
- Forensic reconstruction: Reconstructing analyses may require comparing information across papers and journals, using ordering, off-by-one checks, and cross-study cell-line overlaps.The authors describe this process as a form of archaeology around claims established “as previously shown.”
- Clinical implications: The authors state that sensitive/resistant label reversal in pemetrexed may put patients at risk by giving treatment guidance opposite to the truth if the general approach works.This consequence is presented conditionally on the broader approach being valid.
- Assessment of the approach: When applied without the identified errors, predictions from NCI-60 cell lines were no better than chance, according to the authors’ experiments.The authors report reaching an impasse despite communicating their concerns to the original authors.
- Reproducibility practices: The authors argue that complete scripts will eventually be required alongside raw data and cite REMARK guidelines for detailed reporting expectations.They also instituted Sweave-based reports, earlier collaborator discussions, and standardized analysis templates.
Supplement A: Examining doxorubicin in detail
Supplement A documents a report on identifying ties and sensitive/resistant status for samples checked for doxorubicin.
- Supplement contents: Supplement A contains a zipped PDF report describing identification of ties and sensitive/resistant status for doxorubicin samples.The supplement is identified by DOI 10.1214/09-AOAS291SUPPA.
Supplement C: Examining combination therapy
Supplement C contains reports examining testing numbers, clinical blocks and confounding, combination rules, gene lists, ROC curves, and cell-line drug-sensitivity values.
- Supplement contents: Supplement C contains seven PDF reports covering testing numbers, clinical blocks, confounding, combination rules, gene lists, prediction of cyclophosphamide sensitivity, and cell-line checks.The reports include getTestingNumbers, getTestingClinical, checkOldCombinationRule, checkDrugSensitivity, mapGeneLists, predictCytoxanSensitivity, and checkingCellLines.
- Training-cell-line analysis: The combination-therapy supplement includes a report identifying the training cell lines needed to produce the cyclophosphamide gene list.This report is part of Supplement D’s cell-line survey materials.
Supplement E: Examining docetaxel in detail
This supplement reports on docetaxel samples, including tie identification and sensitive/resistant status, and provides Sweave reports with additional data and code.
- The report identifies ties and classifies checked samples as sensitive or resistant to docetaxel.
- All reports are produced in Sweave.
- Additional reports, data, and code are available through the ReproRsch-All supplement.
- Reports on combination therapy are provided in the ReproRsch-Breast supplement.