In a Google Research experiment, European genomic data improved predictions for Japanese participants when the local sample was small
In a Google Research experiment, European genomic data improved predictions for Japanese participants when the local sample was small
On 3 September, researchers varied the sizes of the European and Japanese training samples and evaluated predictions for eight health traits using a separate subset of Biobank Japan data. The value of the external dataset depended on both the number of local genomes and the trait being predicted.
A polygenic score combines the contributions of many DNA variants to predict a biological trait or disease risk. These models are often trained on large datasets that contain particularly high numbers of people of European ancestry. Accuracy can decline when a model is applied to another population because genetic variants may differ in frequency and may have different associations with the same trait.
In a Google Research experiment, researchers compared the European subset of UK Biobank, a British biobank containing genetic and medical data, with Biobank Japan, a Japanese biobank with nearly 200 thousand participants. They varied the sizes of the two training datasets for eight traits, including body mass index, blood pressure, blood measurements, HDL cholesterol, LDL cholesterol, and glucose. They then evaluated every model on the same separate subset of Biobank Japan data, which had not been used for training.
In the first series of analyses, the researchers initially identified predictive genetic variants in the European UK Biobank data. With approximately 5 thousand Japanese samples, adding the European sample improved accuracy because the model had more observations from which to detect weak associations. With 15 thousand Japanese samples or more, the model trained only on Japanese data became more accurate than the model trained on both datasets. For HDL, one measure of cholesterol, adding more than 5 thousand European samples beyond this point reduced accuracy.
The threshold varied by trait. Genetic associations for body mass index were more similar between the two populations, so European data remained useful with 25–40 thousand or more Japanese samples. For HDL, LDL, another measure of cholesterol, and glucose, the highest accuracy required less European data even when the Japanese sample was smaller.
The researchers also compared other methods for incorporating Japanese data. Meta-analysis combined results from the two populations during the variant discovery stage. It was particularly helpful when the Japanese samples were small and the genetic associations for a trait differed between populations. PRS-CSx generated separate predictions for each population and then combined them. This method required more Japanese data to match the accuracy of the best-performing model from the first series.
In this experiment, the external dataset gave the model additional statistical support while local data were limited. As the local sample grew, its data described the target population more accurately. The appropriate balance among transferring external data, expanding the local biobank, and choosing a training method therefore depended on both the size of the local sample and the specific trait.