CDISC WorkflowSDTM · ADaM · TFL · R QC
QC: PASS · 0 differences

HF-1002-CL-101 · CDISC pipeline

Raw EDC to SDTM, ADaM, and TFL, double-programmed in R

A worked CDISC submission pipeline on a fictional Phase 1 study. Every layer is built in Python and independently re-derived in R; the two are compared to zero differences.

Starting from a fictional Phase 1 protocol and its SAP, the pipeline maps the raw EDC into SDTM tabulation datasets, derives the ADaM analysis datasets, and produces the TFL tables. Each layer is built in Python and independently re-programmed in R, then compared cell by cell.

50
27 SDTM + 7 ADaM datasets + 16 TFL tables
18
subjects · 3 cohorts · 19 EDC forms
0
QC differences (228,191 values compared)
1
real defect caught (lossy CSV float parser)

The pipeline

  1. 01Protocol & SAP
  2. 02Raw EDC (19 forms)
  3. 03SDTM · 27 datasets
  4. 04SDTM QC (R)
  5. 05ADaM · 7 datasets
  6. 06ADaM QC (R)
  7. 07TFL · 16 tables
  8. 08TFL QC (R)

Each layer reads the one before it: raw EDC → SDTM → ADaM → TFL. The red steps are the independent R QC — a second implementation that must agree to the cell.

Scope. The data is synthetic, built for this demo. The build follows CDISC SDTMIG 3.4 and ADaMIG 1.3 with a generated define.xml 2.1 and conformance checks at each layer. QC is independent double programming in R (diffdf for datasets, cell-by-cell for tables). The AE/CM/MH coding is a synthetic stand-in, not licensed MedDRA/WHODrug.

Source documents

The pre-specified inputs the whole pipeline is built from.

Source
Clinical Study Protocol
The fictional Phase 1, single-arm, 3-cohort dose-escalation protocol.
HF-1002-CL-101_Protocol_v1.0_06Jun2026.docxOpen
SAP
Statistical Analysis Plan
Source of truth
Analysis sets, endpoints, and the table inventory the TFL layer follows.
SAP_HF-1002-CL-101_v1.0_06Jun2026.docxOpen
Shells
Mock TFL Shells
The empty table shells the populated TFL package is built to match.
HF-1002-CL-101_Mock_Shells.docxOpen

SDTM tabulation datasets

The 27 SDTM datasets, built from the 19 raw EDC forms to SDTMIG 3.4. A single metadata spec drives the XPT labels/types, the define.xml 2.1, and the dataset spec, so all three stay consistent.

27
datasets (XPT + CSV)
118,078
values written
0
conformance errors
27/27
datasets pass R QC
Notable derivations. Custom Findings domains ZE (echocardiography) and PE are added where no standard domain exists. AE EPOCH is derived from each event's start date against the subject's elements, not the CRF collection folder (all AEs are logged on one Day-1 form). Baseline flags, study days, and --SEQ follow the documented keys so the R build reproduces them exactly.

Adverse events, subject 1001 (SDTM AE)

SubjectReported termDictionary termStartEpochGrade
1001Catheterization site haematomaCatheter site haematoma2026-06-15TREATMENT1
1001PyrexiaPyrexia2026-06-15TREATMENT1

Verbatim term plus synthetic-coded preferred term, with the derived analysis epoch and CTCAE grade.

Domain inventory

Special Purpose

DomainLabelRowsVars
DMDemographics1824
SESubject Elements7010
SVSubject Visits25510

Interventions

DomainLabelRowsVars
CMConcomitant Medications11118
EXExposure1816

Events

DomainLabelRowsVars
AEAdverse Events5722
CEClinical Events2112
DSDisposition11813
MHMedical History7111

Findings

DomainLabelRowsVars
EGECG Test Results72617
FTFunctional Tests20118
ISImmunogenicity Specimen Assessments57318
LBLaboratory Test Results213124
PEPhysical Examination (Custom)13813
QSQuestionnaires23516
VSVital Signs99118
ZEEchocardiography (Custom)26819

Trial Design

DomainLabelRowsVars
TATrial Arms1210
TETrial Elements46
TITrial Inclusion/Exclusion Criteria145
TSTrial Summary167
TVTrial Visits159

Relationship

DomainLabelRowsVars
RELRECRelated Records87
SUPPAESupplemental Qualifiers for AE5710
SUPPDSSupplemental Qualifiers for DS210
SUPPEXSupplemental Qualifiers for EX10810
SUPPPESupplemental Qualifiers for PE2010

ADaM analysis datasets

The 7 ADaM datasets, derived from SDTM to ADaMIG 1.3. ADSL is the subject-level hub carrying treatment and population flags; the BDS datasets add analysis values with baseline, change, and visit windowing.

DatasetLabelClassRows
ADSLSubject-Level Analysis DatasetAdsl18
ADAEAdverse Events Analysis DatasetOccurrence Data Structure57
ADEXExposure Analysis DatasetBasic Data Structure54
ADLBLaboratory Analysis DatasetBasic Data Structure1834
ADVSVital Signs Analysis DatasetBasic Data Structure830
ADIMGImaging Analysis DatasetBasic Data Structure67
ADEFFEfficacy Analysis DatasetBasic Data Structure335
Analysis structure. Population flags ENRLFL / SAFFL / FASFL / DLTFL follow the SAP analysis-set definitions. The shared BDS engine derives AVISIT windowing, ABLFL, BASE, CHG/PCHG, and ANL01FL. Dates are stored as SAS numeric days so the R build compares exactly.

Change from baseline in LVEF, subject 1001 (ADIMG, central reader)

Analysis visitAVALBASECHGABLFLANL01FL
Baseline28.928.9YY
Month 332.628.93.7Y
Month 629.028.90.1Y
Month 1228.428.9-0.5Y

Baseline row highlighted. CHG = AVAL − BASE; LVEF improves over the 12-month follow-up. These analysis-ready values feed the efficacy tables directly.

Define & data explorer

Browse every SDTM and ADaM dataset in the browser — the rows (Data) or the variable-level metadata (Define) — or download the submission-format files. The define metadata is rendered from the same build manifest that produces define.xml.

SDTM
SDTM metadata
Browse the define-XML as a themed page, or download the raw define.xml and the SDTM specification workbook.
ADaM
ADaM metadata
Browse the define-XML as a themed page, or download the raw define.xml and the ADaM specification workbook.

Row data loads on demand from the CSV; the metadata view needs no fetch. Downloads above are the XPT-package metadata; per-dataset CSV/XPT downloads sit in the header when a dataset is selected.

Populated tables

All 16 tables from the mock-shell inventory, generated directly from the QC'd ADaM datasets (RTF, fixed-width text, and CSV). The text shown here is the generator's actual output. Pick a table.

Table 14.1.1
Demographics and Baseline Characteristics
Analysis Population: Safety Set

                                            Cohort 1 (Low)    Cohort 2 (Mid)   Cohort 3 (High)        Total       
                                                (N=6)             (N=6)             (N=6)             (N=18)      
------------------------------------------------------------------------------------------------------------------
Age (years)
  n                                               6                 6                 6                 18        
  Mean (SD)                                  58.7 (14.19)      60.0 (10.73)      60.5 (4.89)       59.7 (10.04)   
  Median                                         54.5              59.0              61.0              59.0       
  Min, Max                                    43.0, 76.0        43.0, 75.0        53.0, 66.0        43.0, 76.0    
Age Group, n (%)
  <65 years                                    4 (66.7)          4 (66.7)          4 (66.7)         12 (66.7)     
  >=65 years                                   2 (33.3)          2 (33.3)          2 (33.3)          6 (33.3)     
  >=75 years                                   2 (33.3)          1 (16.7)          0 (0.0)           3 (16.7)     
Sex, n (%)
  Male                                         4 (66.7)          3 (50.0)          4 (66.7)         11 (61.1)     
  Female                                       2 (33.3)          3 (50.0)          2 (33.3)          7 (38.9)     
Race, n (%)
  White                                        3 (50.0)          3 (50.0)          3 (50.0)          9 (50.0)     
  Black Or African American                    2 (33.3)          2 (33.3)          1 (16.7)          5 (27.8)     
  Asian                                        1 (16.7)          1 (16.7)          2 (33.3)          4 (22.2)     
Ethnicity, n (%)
  Hispanic Or Latino                           3 (50.0)          1 (16.7)          1 (16.7)          5 (27.8)     
  Not Hispanic Or Latino                       3 (50.0)          5 (83.3)          5 (83.3)         13 (72.2)     
Height (cm)
  n                                               6                 6                 6                 18        
  Mean (SD)                                 171.2 (11.18)      179.1 (4.81)      171.8 (7.89)      174.0 (8.68)   
  Median                                        172.3             179.7             169.9             175.4       
  Min, Max                                   159.1, 183.9      171.0, 184.3      163.8, 186.5      159.1, 186.5   
Weight (kg)
  n                                               6                 6                 6                 18        
  Mean (SD)                                  78.8 (13.61)      76.1 (9.55)       90.3 (8.73)       81.7 (12.00)   
  Median                                         74.9              76.3              91.8              78.0       
  Min, Max                                   62.7, 102.9        63.0, 91.5       76.8, 100.4       62.7, 102.9    
BMI (kg/m2)
  n                                               6                 6                 6                 18        
  Mean (SD)                                  26.9 (3.51)       23.9 (4.21)       30.6 (2.39)       27.1 (4.30)    
  Median                                         27.8              24.0              30.5              28.2       
  Min, Max                                    22.6, 30.4        19.1, 31.3        27.4, 33.3        19.1, 33.3    
------------------------------------------------------------------------------------------------------------------
Source: ADSL
[a] SD = Standard Deviation; BMI = Body Mass Index.
[a] Percentages are based on N in the column heading.

Columns are the three dose cohorts plus Total; populations follow the SAP. Every cell was independently reproduced in R (see QC).

Independent QC, in R

Every layer was re-programmed from scratch in R — reading the same inputs but sharing no code — and compared against the Python production. Datasets are compared with diffdf (the PROC COMPARE equivalent); the tables are compared cell by cell. Result across all three layers: zero differences.

27/27
SDTM datasets pass
7/7
ADaM datasets pass
16/16
TFL tables pass (cell-by-cell)
1
real defect caught and fixed

Coverage

LayerComparison methodScopeResult
SDTMdiffdf dataset compare27 datasets27/27 PASS
ADaMdiffdf dataset compare7 datasets7/7 PASS
TFLcell-by-cell compare16 tables · 1,365 cells16/16 PASS

The defect double-programming caught

Lossy float parsing. The first TFL QC run flagged two table cells. Root cause: pandas.read_csv's default float parser is lossy (off by ~1 ULP) and parsed 19.799999999999997 as 19.800000000000001, diverging from R's correct parser exactly at a rounding boundary — so the Python tables rounded two cells the wrong way. Fixed by reading with float_precision="round_trip". A single-implementation build, and a structural check, would both have missed this; only a second independent implementation disagreeing surfaced it.

Each QC harness was validated with a negative control — injecting a wrong value or dropping a row makes the comparison fail — so an all-pass result is meaningful, not trivial. Reports: sdtm_out/qc/, adam_out/qc/, tfl_out/tfl_qc/.

Demonstration. Protocol HF-1002-CL-101 (fictional) · Phase 1 gene therapy · synthetic data · SDTMIG 3.4 / ADaMIG 1.3 · not for clinical or submission use