I absorb the engineering in biobank-scale research, so the research is what’s left.

Everything below the line is also me. It just didn’t reach significance.

I am a bioinformatics scientist in the Precision Health Informatics Section at NHGRI. Most of what I produce is a library, a pipeline, or a query layer, and the analyses on top of it. It is infrastructure that lets a research question be asked at biobank scale, and asked again later.

The larger share of that is engineering: Python, SQL, and workflow infrastructure on GCP and AWS, pointed at All of Us, UK Biobank, and the electronic health record and genomic data underneath them. A smaller but regular part of the work is the analysis itself, association studies and the statistical analyses common to large cohorts. The two are difficult to separate in practice: what the analysis needs is usually what determines how the pipeline gets built.

A separate line of work is sequencing analysis: alignment and quality control, bulk and single-cell RNA-seq, and the downstream steps of differential expression, gene set enrichment, and pseudotime. This one runs through everything I have done: graduate training, then industry, and now collaborations and consulting.

I came to this from the other end of the bench. Wet-lab work across microbiology, immunology, and molecular biology, in both academic and industry settings. Much of it ran under GLP and GMP, where process control and documentation are held to the same standard as the science. A Lean Six Sigma Black Belt along the way gave me a structured way to work through a problem and improve the process behind it, which has stayed useful in every kind of work I have done since.

I no longer run those assays, but knowing what data went through before it reaches me is often useful. That background is also why I treat a pipeline as something readable, auditable, and reproducible, not just something that runs.

Languages
Python, SQL, R, Bash
Cloud
GCP, AWS, BigQuery
Workflow
dsub, Docker
Version control & CI
Git, GitHub Actions
Data
OMOP / EHR, genomic variant data, phecodes
Sequencing
End-to-end bulk and single-cell sequencing analysis
Certification
Lean Six Sigma Black Belt

Most of what I build is engineering: the library, the pipeline, the query layer that lets a question be asked at biobank scale. The tools come first here, then what happens when they meet a real analysis.

PheTK

A fast, efficient, and resource-friendly Python library for PheWAS analysis

Five modules that compose into one workflow: clinvar selects variants by clinical significance, cohort builds the study population, phecode maps diagnosis codes to phenotypes, phewas runs the regressions, plot draws the result.

Citing papers per year citing papers per year 2 2024 7 2025 16 2026 to Aug preprint published · Bioinformatics, Dec 2024

Running a PheWAS at biobank scale usually means assembling several tools and writing the glue between them. PheTK covers the whole path: genetic cohort, covariates, ICD-to-phecode mapping and phecode counts, the regressions, the plot. On the All of Us Researcher Workbench it covers it directly, with functions that already know where the data sits and what shape it is in. The queries, the joins, the reshaping between steps all stay inside the package, so what is left to the researcher is the analysis: choose a variant, choose covariates, choose a model, read the result. Both phecode 1.2 and phecodeX 1.0 are supported, and the model can be logistic or Cox, with or without Firth penalization.

Underneath, it is built for the scale that arrived since most PheWAS packages were written, when a large cohort meant tens of thousands of people. It has been tested to a million participants, roughly twice the largest cohort that currently has the health record data a PheWAS needs (UK Biobank, 500,000; All of Us, ~481,000 with linked EHR). Speed and footprint come out of the same choices: efficient dependencies and data formats, implementations written to fit the problem, and parallelization applied where the method actually benefits from it. On an end-to-end workflow it is substantially faster than the established alternatives it has been benchmarked against.

The All of Us support is convenience, not dependency. The PheWAS function itself is platform-independent and runs on modest resources, a smaller cluster or a laptop, where it finishes rather than running out of memory. And on the platforms the data cannot leave, compute is billed by the hour, so a lighter run is a cheaper one, which keeps the same analysis affordable for a single researcher rather than only a well-funded group. It has been cited by twenty-five papers since release, including in Nature, Nature Genetics, and the Journal of Clinical Investigation.

Countries of citing papers: shaded by any author, dotted by first author United Kingdom 4 papers · 3 institutions Germany 2 papers · 2 institutions United States 15 papers · 14 institutions China 2 papers · 2 institutions India 1 paper · 1 institution Australia Austria Brazil Canada China Croatia Estonia France Germany India Ireland Italy Japan Latvia Netherlands Nigeria Norway Poland Switzerland Taiwan United Kingdom United States Zambia
Shading marks every country where an author on a citing paper is based. Dots mark first authors only, so those counts add up rather than overlapping; hover a dot for its counts, or a shaded country for its name. Citation counts from Google Scholar; affiliations from PubMed and the preprint servers, retrieved September 2026.

Cloud analysis pipelines

Engineering, NHGRI

Planning and running large analyses so they finish inside a budget a researcher can actually commit to.

A library that is inexpensive to run only helps if the run around it is planned. Collaborators have used PheTK for thousands of PheWAS analyses against biobank data at a few hundred dollars of compute, partly because the library is efficient, and partly because the work is sequenced: plan the run, test at small scale, scale up in stages, then commit to the full analysis, so the expensive mistakes happen while they are still cheap.

The rest is machine sizing, containers so a run is reproducible rather than reconstructed, and knowing which parts of a run are worth paying to parallelize. Across these projects it adds up to more than $50,000 of compute that was not spent.

Collaboration

NHGRI
Papers, field, journal and author position one dot per author · the filled dot is me · counted where the list is too long to draw yearfirst authorfieldpaperjournal authors 2026CambridgeProteomicsCardiovascular proteome medRxiv7 / 34 NIEHSImmunogeneticsHLA across ancestries medRxiv OxfordDigital healthWearable activity data medRxiv 2025NHGRIGeneticsCystic fibrosis carriers JAMA Intern Med VanderbiltOncologyProstate cancer, cardiac risk The Prostate ArizonaGeneticsAPOE across the phenome eBioMedicine NHGRIInfectious diseaseRespiratory virus phenotyping Sci Rep 2024NHGRIMethodsPheTK Bioinformatics NHGRIPharmacologyAntidepressants, hyponatremia Clin Pharmacol Ther NHGRIPharmacologyAntihypertensive prescribing Clin Pharmacol Ther NCIImmunologyGerminal centre immunology Nat Immunol NHGRIMethodsAll of Us vs UK Biobank JAMIA ManchesterGeneticsType 2 diabetes heterogeneity Nature235 / 363 2023NHGRIEpidemiologySmoking associations JAMIA S. DakotaPharmacologyAntibiotic liver injury Clin Pharmacol Ther MalawiImmunologyCerebral malaria Malar J 10 distinct 10 distinct 16, 1 first-authored 11 distinct
Author lists verified against PubMed, September 2026.

A selection, newest first. The full list is on Google Scholar.

  1. Diseases common in persons with cystic fibrosis among CFTR heterozygotes

    Zeng C, Han ST, Cassini TA, Raraigh KS, Tran TC, Yang J, Cutting GR, Denny JC

    JAMA Internal Medicine 185(8):1014–1024, 2025doi:10.1001/jamainternmed.2025.1853

  2. Phenome-wide association of APOE alleles in the All of Us Research Program

    Khajouei E, Ghisays V, Piras IS, Martinez KL, Vicenti AT, Naymik M, Ngo P, Tran TC, et al.

    EBioMedicine 117:105768, 2025doi:10.1016/j.ebiom.2025.105768

  3. PheWAS analysis on large-scale biobank data with PheTK

    Tran TC, Schlueter DJ, Zeng C, Mo H, Carroll RJ, Denny JC

    Bioinformatics 41(1), 2025doi:10.1093/bioinformatics/btae719First author

  4. Gα13 restricts nutrient driven proliferation in mucosal germinal centers

    Nguyen HT, Li M, Vadakath R, Henke KA, Tran TC, et al.

    Nature Immunology 25(9):1718–1730, 2024doi:10.1038/s41590-024-01910-0

  5. Comparison of phenomic profiles in the All of Us Research Program against the US general population and the UK Biobank

    Zeng C, Schlueter DJ, Tran TC, Babbar A, Cassini T, Bastarache LA, Denny JC

    Journal of the American Medical Informatics Association 31(4):846–854, 2024doi:10.1093/jamia/ocad260

  6. Drug-induced liver injury with commonly used antibiotics in the All of Us Research Program

    Gu S, Rajendiran G, Forest K, Tran TC, Denny JC, Larson EA, Wilke RA

    Clinical Pharmacology & Therapeutics 114(2):404–412, 2023doi:10.1002/cpt.2930

  7. The ability of interleukin-10 to negate haemozoin-related pro-inflammatory effects has the potential to restore impaired macrophage function associated with malaria infection

    Tembo D, Harawa V, Tran TC, Afran L, Molyneux ME, Taylor TE, et al.

    Malaria Journal 22(1):125, 2023doi:10.1186/s12936-023-04539-w

  1. 2022–present

    Bioinformatics Scientist

    Precision Health Informatics Section, NHGRI, NIH

  2. 2020–present

    Bioinformatics and data science consultant

  3. 2019–2020

    Biologist

    Microbiome Core, NIAID, NIH

  4. 2018–2019

    Associate Scientist II/III

    NGS and bioinformatics core, BioReliance / MilliporeSigma

  5. Education

    MS Analytics, Georgia Institute of Technology, USA

    MS Microbiology and Immunology, Cornell University, USA

    MS Biotechnology, University of East Anglia, UK

    BS Biotechnology, University of Sciences, Vietnam

The full version is available on request at tam.c.tran@outlook.com.

Photographs taken outside work, which turns out to need much the same patience as the day job.

View the gallery