PONG 2.0 reads immune genetics from cheap DNA chip data already in biobank samples with comparable accuracy across African, South Asian, East Asian, European, and Admixed American samples.
A new free R package called PONG 2.0 reads a hidden layer of immune genetics out of the cheap DNA-chip data already sitting in millions of biobank samples, and it does so with comparable accuracy across African, South Asian, East Asian, European, and Admixed American samples. The package, published this year in Human Molecular Genetics by its authors, imputes high-resolution KIR genotypes, the on/off switches that natural killer cells use to read neighboring cells, directly from SNP-array data that researchers have been generating for over a decade.
The release matters because KIR has long been a blind spot in genome-wide association studies. The KIR gene cluster on chromosome 19 is so copy-number-variable that standard genotyping chips cannot read it, and the HLA class I ligands it reads, the molecular ID badges every cell wears so the immune system can recognize it, are themselves the most polymorphic region of the human genome. Together, that means GWAS catalogs, built around SNP arrays, have effectively skipped a layer of immune variation that drives transplant matching, infection control, pregnancy, autoimmune disease, and cancer outcomes.
PONG 2.0 is a statistical inference engine that turns the sparser SNP-array readout into a detailed KIR call. The authors trained it on a multi-ancestry subset of the 1000 Genomes Project: 187 European, 93 Admixed American, 102 South Asian, 102 African, and 102 East Asian samples. Reported training-set accuracy landed between 92 and 99 percent. On a held-out cohort of 267 samples typed by targeted sequencing, per-locus concordance ranged from 92.1 to 97.7 percent. At population scale, the package recovered KIR allele frequencies from more than 8,000 individuals with R² up to 0.999 and median deviation between 0.6 and 6.8 percent.
The per-locus accuracy floor (92.1 percent) sits inside the range researchers accept for HLA imputation from SNP arrays, and the population-level R² values mean association signals will largely survive the imputation step. The multi-ancestry training and validation are the structural news: earlier KIR imputation tools were calibrated almost entirely on European samples, so non-European biobanks could not get reliable calls without re-typing.
The cost of asking a KIR question in a biobank has just collapsed. Take transplant rejection. Donor-recipient KIR-HLA matching is already used in haploidentical bone-marrow transplant protocols, and retrospective studies in non-European cohorts have been limited by the cost of high-resolution KIR typing. A multi-ancestry GWAS of NK-cell activity in graft-versus-host disease, run on existing biobank arrays, is now a doctoral student and a compute job rather than a custom sequencing project. The same pattern applies to large-scale questions about KIR3DL1/HLA-B combinations in HIV control, KIR2DS1 in pregnancy disorders, and the KIR region's role in type 1 diabetes and psoriasis, all of which have been chronically underpowered because the genotyping step was too expensive to scale.
The package ships as an open-source R release with pre-trained models at github.com/NormanLabUCD/PONG2. The team frames it as research-grade imputation, not a clinical assay, and the accuracy range varies by locus and by population representation within the 1000 Genomes reference. Imputation quality is also a function of the SNP-array platform, and labs running older chips should expect the lower end of the concordance range. Independent validation here means held out from the training pipeline, not an external wet-lab replication study, so clinical translation, using PONG 2.0 calls to guide a transplant decision, is not what this release supports.
The bottleneck that kept KIR out of GWAS catalogs is now a software install. The downstream question is whether the field uses it before the next wave of biobank-scale whole-genome sequencing makes imputation moot, or whether the multi-ancestry reference turns out to be a permanent public good either way.