All projects
FLAGSHIP · SDTM ENGINE

SDTM Engine

One declarative YAML spec per domain turns a raw EDC export into a CDISC SDTM package: data, mapping spec, and Define-XML from a single source of truth, judged against SDTMIG 3.4 + CDISC CT, never a vendor template. Every mapped value traces to a real collected field; nothing is fabricated. Proven on a real multinational study, it builds the analysis-critical science correctly, and is honest about the supportive scaffolding that is still mechanical to finish.

Launch live demo Runs in your browser
Run on a real study, not just demos

This demo uses synthetic data so you can run it freely. The same engine has been run end to end on a real cardiovascular study: raw EDC export to SDTM, 12 domains, 26,973 records, validated against the external CDISC CORE engine, with zero source-trace gaps. Every value carried a path back to its raw field. Where the source data had problems, a blank required field, a controlled-terminology extension, the engine reported them rather than filling them in. It does not fabricate a value to clear a check. That run is the proof the pipeline holds up outside a sandbox; the demo below shows how it works.

What it does

Building the CDISC SDTM datasets from a raw EDC export is the clinical deliverable most sponsors outsource. This engine does it from one declarative YAML spec per domain. The spec drives both the data build and the rendered mapping spec, so the data and its documentation cannot drift, and conformance is judged against bundled SDTMIG + CDISC controlled terminology rather than a vendor package. Distinct from the CDISC Submission Pipeline: this is the spec-driven mapping engine, traceable to the raw field, not double-programming reconciliation.

How it works

CDISC SDTMSDTMIG 3.4 Controlled terminologyDefine-XMLClick-traceability

Why SDTM and ADaM are not the same problem

SDTM is standardization. Raw data mapped into a published structure with fixed domains, variables, and controlled terminology. There is a correct answer the standard defines, which is why this layer can be built to a button: given the data and the spec, the output is deterministic, reproducible, and conformance-checkable.

ADaM is analysis. There is no published "correct" ADaM, because ADaM is constructed to serve a specific statistical analysis plan. What counts as baseline, how treatment-emergence is defined, which population, which windows, how missing visits are handled, these are judgment calls that come from the protocol and the SAP and differ for every study. That layer cannot be a button, and anyone who tells you it can has not built one.

So the two layers are built differently, and the showcase keeps them as separate demos. SDTM is push-button: this engine aims for clean, traceable, conformance-checkable output. ADaM is the job of the Statistical Programming Pipeline demo, which does the mechanical work but surfaces the judgment calls and traces every derivation, so a statistician makes and owns the decisions instead of a black box making them silently. The goal is not to remove the statistician. It is to make the work faster and the decisions auditable.

Synthetic data here is a choice, not a limit: it lets you run the demo freely and inspect every derivation. The engine behind it has been validated on a real study against CDISC CORE.

A from-scratch demo built from a spec-driven SDTM engine. Synthetic seeded study and raw EDC; every value is computed in the browser with an in-page self-test that reconciles the build, file://-safe with no network calls. Takes ?tab= (overview, build, conformance, trace, compare) and ?still=1. No real subjects, no sponsor data.

Screenshot of the SDTM Engine demo: a three-panel build-and-trace view linking a raw EDC field to a YAML spec op to the built SDTM cell Static preview - click to open the live, interactive demo