Data Sources & Provenance

Overview

This vignette documents all data sources used in the coMMpass analysis pipeline, their access tiers, and the distinction between pipeline data (real patient data from GDC) and synthetic test data (generated by example_data()).

MMRF CoMMpass Study

The Multiple Myeloma Research Foundation (MMRF) Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile (CoMMpass) study is a longitudinal observational study of ~1,143 newly diagnosed multiple myeloma patients.

Citation

Keats JJ, et al. Interim Analysis Of The Mmrf CoMMpass Trial, a Longitudinal Study In Multiple Myeloma Relating Clinical Outcomes to Genomic and Immunophenotypic Profiles. Blood. 2013;122(21):532. doi:10.1182/blood.V122.21.532.532

Data at a Glance

Data Access Tiers

Tier Data Access Used by pipeline
Open Access RNA-seq counts (STAR), clinical metadata, treatment records GDC Data Portal, no login required Yes
Open Access (S3) RNA-seq files s3://gdc-mmrf-commpass-phs000748-2-open/ Yes
MMRF Gateway FISH/cytogenetics, PFS, treatment response Free registration (pending access) No — 12 targets blocked
Controlled Access WGS, WES, protected clinical dbGaP application (phs000748) No — requires institutional IRB

Pipeline Data Sources (Real)

The targets pipeline downloads real patient data from GDC:

Data Source Function
RNA-seq SummarizedExperiment GDC STAR-Counts download_rnaseq_data() via TCGAbiolinks::GDCdownload()
Clinical metadata (demographics, ISS, vital status) GDC clinical endpoint download_clinical_data() via TCGAbiolinks::GDCquery_clinic()
Treatment records (7,184 records, 994 patients) GDC API download_clinical_data() via GDC REST API
Biospecimen metadata GDC biospecimen endpoint download_clinical_data() via TCGAbiolinks::GDCquery_clinic()
MSigDB gene sets (Hallmark + KEGG) MSigDB Pre-bundled parquet in inst/extdata/msigdb/

The pipeline runs with a configurable sample_limit parameter (default 200 in local builds; CI uses 20). All vignettes load pre-computed results from the targets store.

GDC Documentation

Synthetic Test Data (example_data())

The example_data() function returns small synthetic datasets for testing and documentation examples. These are randomly generated with set.seed(42) and contain no real patient data.

Dataset Size Description
rnaseq_se 50 genes x 20 samples SummarizedExperiment with simulated counts
clinical 20 patients Simulated demographics and outcomes
cytogenetic 20 patients Simulated FISH markers and risk groups
treatment ~40 rows Simulated treatment lines (1-3 per patient)

Stored in inst/extdata/example/*.rds and created by create_all_example_data().

The two data tracks do not mix: pipeline vignettes use real GDC data via tar_read(); unit tests and ?example_data examples use synthetic data only.

MSigDB Gene Sets

Pre-bundled MSigDB Hallmark gene sets are stored as parquet files in inst/extdata/msigdb/. These are public reference gene sets used for pathway enrichment analysis.

Bundled Files

Citation

Liberzon A, et al. The Molecular Signatures Database (MSigDB) Hallmark Gene Set Collection. Cell Systems. 2015;1(6):417-425. doi:10.1016/j.cels.2015.12.004

Recent Changes

Recent project commits with lines added, files changed, and change categories.

Reproducibility

Session Info (click to expand)
Show code
sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 26.04 LTS
#> 
#> Matrix products: default
#> BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.32.so;  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=en_US.UTF-8       LC_NUMERIC=C              
#>  [3] LC_TIME=en_US.UTF-8        LC_COLLATE=en_US.UTF-8    
#>  [5] LC_MONETARY=en_US.UTF-8    LC_MESSAGES=en_US.UTF-8   
#>  [7] LC_PAPER=en_US.UTF-8       LC_NAME=C                 
#>  [9] LC_ADDRESS=C               LC_TELEPHONE=C            
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C       
#> 
#> time zone: Etc/UTC
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats     graphics  grDevices utils     datasets  methods   base     
#> 
#> other attached packages:
#> [1] targets_1.12.0
#> 
#> loaded via a namespace (and not attached):
#>  [1] vctrs_0.7.3       cli_3.6.6         knitr_1.51        rlang_1.3.0      
#>  [5] xfun_0.60         otel_0.2.0        processx_3.9.0    jsonlite_2.0.0   
#>  [9] data.table_1.18.4 glue_1.8.1        prettyunits_1.2.0 DT_0.34.0        
#> [13] buildtools_1.0.0  backports_1.5.1   htmltools_0.5.9   maketools_1.3.2  
#> [17] sys_3.4.3         ps_1.9.3          sass_0.4.10       rmarkdown_2.31   
#> [21] jquerylib_0.1.4   crosstalk_1.2.2   tibble_3.3.1      evaluate_1.0.5   
#> [25] base64url_1.4     fastmap_1.2.0     yaml_2.3.12       lifecycle_1.0.5  
#> [29] compiler_4.6.1    codetools_0.2-20  igraph_2.3.3      htmlwidgets_1.6.4
#> [33] pkgconfig_2.0.3   digest_0.6.39     R6_2.6.1          tidyselect_1.2.1 
#> [37] pillar_1.11.1     callr_3.8.0       magrittr_2.0.5    bslib_0.11.0     
#> [41] withr_3.0.3       tools_4.6.1       secretbase_1.3.0  cachem_1.1.0