I. DEPARTMENT INFORMATION
| Job Description Summary: |
The Biostatistics Center (
BSC) of the Milken Institute School of Public Health is an off-campus research facility of The George Washington University located in Rockville, Maryland. The
BSC serves as the coordinating center for large scale multi-center clinical trials and epidemiological studies funded by federal agencies including the National Institutes of Health. The
BSC is a leader in the statistical coordination of major medical research programs of national and international scope. Visit our website at:
www.bsc.gwu.edu.
We are seeking a
Research Scientist specializing in
Machine Learning, Data Science, Data Harmonization, and Synthetic Data Generation to lead the integration, standardization, and privacy-preserving algorithmic modeling of complex, multi-site datasets. In this role, you will bridge the gap between complex data infrastructure, cutting-edge machine learning, and synthetic data generation-building automated transformation pipelines, harmonizing disparate clinical/observational data structures (e.g.,
OMOP
CDM,
FHIR), and generating high-fidelity synthetic datasets to accelerate secure collaborative research without compromising data privacy.
Experience with high-dimensional multi-omics data analysis and integration is highly desirable.
Key Responsibilities:
Synthetic Data Generation & Privacy Engineering
Generative Modeling: Design, train, and validate generative machine learning models (GANs, VAEs, diffusion models, and LLM-based tabular synthesizers) to generate high-fidelity synthetic tabular, longitudinal, and multi-omic datasets. Privacy Assurance & Risk Assessment: Implement rigorous privacy-preserving methodologies (differential privacy, membership inference attack testing, re-identification risk metrics) to guarantee synthetic datasets meet strict governance and compliance standards. Utility & Fidelity Evaluation: Establish automated benchmark suites comparing distributional fidelity, feature correlations, cross-sectional/longitudinal validity, and downstream task performance between real and synthetic cohorts.
Data Harmonization & Infrastructure
ETL & Data Standardization: Design, execute, and maintain scalable ETL/ELT pipelines to map complex, longitudinal datasets from multi-center registries and disparate source systems into common data models (e.g., OMOP CDM, PCORnet). Ontology Mapping & Quality Control: Implement vocabulary mappings (SNOMED, RxNorm, LOINC, ICD-10) and robust data quality frameworks to handle missingness, inconsistent formatting, and cross-site measurement variation across clinical and molecular datasets. Registry & Pipeline Integration: Architect automated validation workflows ensuring privacy preservation, site blinding, and reproducible data harmonization across large-scale cohort studies.
Machine Learning, Omics & Statistical Modeling
Multi-Omics Data Integration: Analyze, harmonize, and model high-dimensional biological data layers (e.g., genomics, transcriptomics, metabolomics, proteomics) alongside EHR and clinical registry data. Algorithmic Modeling: Develop, evaluate, and deploy advanced machine learning models (gradient boosted trees, neural architectures, ensemble learning) and semi-parametric causal inference techniques on high-dimensional harmonized and synthetic data. Longitudinal & Trajectory Analysis: Build statistical and ML frameworks for analyzing longitudinal trajectories, biological aging markers, and time-varying exposures. Model Explainability & Transportability: Conduct rigorous cross-validation, sensitivity analyses, and transportability evaluations to ensure model fairness, robustness, and biological/clinical validity when trained on real vs. synthetic inputs.
Collaborative Research & Leadership
Interdisciplinary Collaboration: Work closely with biostatisticians, bioinformaticians, privacy officers, software engineers, and domain experts to translate scientific hypotheses into secure, end-to-end data products. Methodological Dissemination: Lead the drafting of peer-reviewed scientific publications, technical whitepapers, and method documentation regarding privacy metrics, synthetic data utility, and multi-omic modeling. Pipeline Open-Sourcing & Mentorship: Contribute to reproducible codebases (R, Python, SQL) and mentor junior data engineers and graduate research assistants.
Performs other work duties as assigned. The omission of specific duties does not preclude the supervisor from assigning tasks logically related to the position. |
| Minimum Qualifications: |
Qualified candidates will hold a Master's degree plus 5 years of experience or a PhD and 2 years of experience in a related discipline to include at least 2 years of research and/or college level teaching in a field basic to the work to be performed. Degree must be conferred by the start date of the position. |
| Additional Required Licenses/Certifications/Posting Specific Minimum Qualifications: |
|
| Preferred Qualifications: |
* Multi-Omics Expertise: Strongly Desirable: Hands-on experience processing, harmonizing, and analyzing high-throughput omics datasets (e.g., genomic variants,
RNA-seq, untargeted metabolomics, panel proteomics) using standard bioinformatics pipelines (e.g., Bioconductor,
PLINK, Nextflow/Snakemake) and integrating them with clinical
EHR/registry frameworks.
* Experience evaluating the "train on synthetic, test on real" (
TSTR) paradigm across diverse biomedical, multi-omic, or observational datasets.
* Experience working with federal funding frameworks (
NIH,
NSF,
PCORI) or multi-center research coordinating centers (BRCCs).
* Familiarity with privacy-preserving technologies (federated learning, differential privacy, secure enclave research environments).
* Knowledge of containerized deployment (Docker, Kubernetes) and version-controlled workflow tools (Git).
* Potential to lead first-author publications in peer-reviewed journals or top-tier machine learning/data science/bioinformatics conferences (e.g., NeurIPS,
KDD,
JAMIA, Biometrics, Bioinformatics). |
| Hiring Range |
$37,949.02 - $61,727.22 (at 57% effort) |
| GW Staff Approach to Pay |
How is pay for new employees determined at GW? |
Healthcare Benefits
GW offers a comprehensive benefit package that includes medical, dental, vision, life & disability insurance, time off & leave, retirement savings, tuition, well-being and various voluntary benefits. For program details and eligibility, please visit https://hr.gwu.edu/benefits-programs.
II. POSITION INFORMATION
| Campus Location: |
Rockville, Maryland |
| College/School/Department: |
The Biostatistics Center |
| Family |
Research and Labs |
| Sub-Family |
Field Research |
| Stream |
Individual Contributor |
| Level |
Level 4 |
| Full-Time/Part-Time: |
Part-Time |
| Hours Per Week: |
20 |
| Work Schedule: |
Monday - Friday, 9am - 1pm |
| Will this job require the employee to work on site? |
Yes |
| Employee Onsite Status |
Hybrid |
| Telework: |
|
| Required Background Check |
Criminal History Screening, Education/Degree/Certifications Verification, Social Security Number Trace, and Sex Offender Registry Search |
| Special Instructions to Applicants: |
Employer will not sponsor for employment Visa status |
| Internal Applicants Only? |
No |
| Posting Number: |
R002466 |
| Job Open Date: |
09/28/2026 |
| Job Close Date: |
|
| Background Screening |
Successful Completion of a Background Screening will be required as a condition of hire. |
| EEO Statement: |
The university is an Equal Employment Opportunity employer that does not unlawfully discriminate in any of its programs or activities on the basis of race, color, religion, sex, national origin, age, disability, veteran status, sexual orientation, gender identity or expression, or on any other basis prohibited by applicable law. |
|