National COVID Cohort Collaborative (N3C): UVA Researchers

The National COVID Cohort Collaborative (N3C) is an NIH-NCATS led project supporting collaborative analytics of COVID data. N3C has built a centralized national data resource that scientists can use to study COVID-19 and identify potential treatments. More than 85 healthcare institutions have contributed over 22 million patient records. The data cannot be removed from the enclave and must be remotely analyzed by researchers who have received the appropriate approvals. UVA has both a Data Transfer Agreement (DTA) and IRB protocol to contribute data as well as a Data Use Agreement (DUA) and IRB approval letter for UVA researcher use of level 3 (limited dataset) N3C data. The SDS Health Informatics team (previously iTHRIV informatics team) has been very involved in developing and using this resource as is evidenced by the below UVA N3C Collaborations map (updated May 25th, 2023, source: National COVID-19 Cohort Collaborative). The informatics team is contracted to provide "Logic Liaison" support to all N3C researchers. Support is provided in the form of reusable Logic Liaison code templates and also research team support at N3C office hours.

ABOUT THE DATA:

The dataset includes medical record information from participating sites about patients who belong to one of the following COVID-19 categories: "lab-confirmed positive," "lab-confirmed negative," "suspected positive," and "possible positive." The current patient cohort selects for the positive cases and then matches them with local lab-confirmed negative patients based on age, sex, race and ethnicity. The positive to negative matching ratio is 1:2, so about 1/3 of the patients in the enclave are COVID-19 positive. The cohort criteria is described in detail in the N3C phenotype on GitHub HERE.

Structured medical record facts (demographics, diagnoses, lab tests, medications, procedures, visit types, etc.) from the cases and controls are sent by participating sites and deposited in the enclave for analysis. The data does not include any personal identifiers besides dates and zip codes which require special approval from a Data Access Committee to access. Data access and analytic resources are free to use in the enclave and are not capped per investigator. Data cannot be exported. Research results can only be represented in aggregate (minimum cell size of 20). Reports of results are reviewed prior to being moved to a download area to ensure no patient or site identifiers are exposed and that the data has been properly aggregated.

GUIDANCE FOR UVA RESEARCHERS CONDUCTING RESEARCH USING N3C DATA:

UVA researchers are encouraged to propose COVID research to be done using this data set and analytics platform. You will need to have Human Subjects Training (as recommended by your institution) in order to use the data. Reference this page for the appropriate training for UVA researchers. Accessing the platform will require NCATS approval of your specific data use request, and you will need to complete NIH security training before accessing the environment. Please be aware that even for users who are well versed in SQL, PySpark/Python, and SparkR/R, navigating and using N3C's Palantir platform as well as acquiring an in-depth understanding of the data tables takes considerable time and effort.

An analysis of this data typically takes several weeks to design/clarify and then a few more weeks to analyze. Feel free to contact Johanna Loomba (jjl4d@uvahealth.org) for support if you are interested in working in N3C for the first time. Support may range from an intro regarding feasibility and onboarding (one hour free consult) all the way up to full service data analysis and biostats support (at cost). A letter from the UVA IRB-HSR is attached which can be reviewed by the study team and attached to a project specific non-human subject research form when applicable.

ADDITIONAL HELPFUL LINKS:

Attribution for collaborative efforts is key to N3C's philosophy of supporting rapid, robust, and reproducible results and is carried out through the Enclave's graph-based tracking and reporting method. The first manuscript that covered the methods for building the Enclave: The National COVID Cohort Collaborative (N3): Rationale, Design, Infrastructure, and Deployment is currently in press at JAMIA with almost 200 authors.

Information about accessing the data is available on the N3C website HERE. The OMOP tables used in N3C are described HERE, and the GitHub repository for OMOP is HERE. Training around use of OMOP data is provided by the Ehden Academy HERE. The current patient cohort is described in the N3C phenotype on GitHub HERE. Please also note that there are specific rules around download of results, publication and citation when you use this data which can be found HERE.

Please note that NCATS is also expanding this enclave to allow researchers to answer questions beyond COVID. More information is HERE.

You may also browse and utilize the N3C Zenodo repository which is a collection of research outputs and information objects produced by and for CTSA Program hubs and that are relevant to the N3C initiative.

Attachments

Filename
Size
Date Modified
N3C_Level_3_Data_UVA_IRB_Determination_12-2022.docx.pdf409.26 KBMay 16, 2026 4:37 AM UTC
N3C for VA partners.pptx2.33 MBMay 16, 2026 4:37 AM UTC
UVA NCATS N3C Data Enclave Institutional Data Use Agreement 21Aug20.pdf240.15 KBMay 16, 2026 4:37 AM UTC