Introduction
In pharmaceutical data science, reproducibility is not just good engineering hygiene — it is part of making an analysis defensible. If someone asks six months later which code, which package versions, and which report produced a result, you need a better answer than “it was in that script somewhere”.
This post is a practical pattern for R-based analysis work in regulated environments: version control, renv, Quarto, and a small amount of discipline around project structure.
Why Reproducibility Matters in Pharma
Regulatory agencies like the U.S. Food and Drug Administration (FDA) and the European Medicines Agency (EMA) require that analytical work be:
- Documented — Clear records of what was done
- Traceable — Ability to follow the analysis from raw data to conclusions ::: {.callout-important} title=“Regulatory Note” Under 21 CFR Part 11 and comparable EMA guidance, “reproducible” means more than running the same script twice. It means you can demonstrate, on demand and later on, exactly which code and which software stack produced a specific result. :::
A practical project structure
The project does not need to be elaborate, but it does need to be legible:
analysis-project/
├── data/
│ ├── raw/
│ └── derived/
├── R/
├── reports/
├── renv.lock
├── analysis.qmd
└── README.md
What matters is the separation of concerns:
- raw data stays untouched
- derived data is generated, not edited by hand
- reusable functions live outside the report
- the report consumes code rather than containing every piece of logic inline
A workflow that holds up in practice
For an individual analysis, my workflow is usually:
- Create the project and commit the skeleton.
- Initialise
renv.
- Write reusable data preparation and analysis functions in
R/.
- Build the final report in Quarto.
- Snapshot the environment once the report renders correctly.
- Commit the code, the lockfile, and the rendered deliverable if the review process requires it.
That sequence is not glamorous, but it prevents the usual failure modes: undocumented package changes, report logic buried in a notebook, and results that cannot be regenerated after a handover.
What to include in the report itself
At minimum, I would include:
- a short description of the analysis objective
- the input data version or extraction date
- the main tables and figures generated from source
- a footer or appendix with
sessionInfo()
- the Git commit hash if the workflow allows it
If the work is high-stakes, add validation notes and explicit assumptions. Reproducibility is stronger when the report explains not only what ran, but what was assumed.
Best Practices
- Use project-relative paths — Never use absolute paths
- Document your environment — Include session info in reports
- Validate data inputs — Check data quality before analysis
- Version control everything — Code, configuration, and documentation
- Separate logic from reporting — Keep reusable functions outside the report body
- Treat the lockfile as part of the analysis — It is not optional metadata
Conclusion
Building reproducible analyses in pharma does require more upfront structure than a one-off script, but that structure pays dividends in quality, compliance, and handover. The stack does not need to be exotic. Git, renv, Quarto, and a clear project layout already get you a long way toward work that can be reviewed and rerun with confidence.
Back to top