menu

arrow_left_alt

Nucleus Labs

Whitepaper

Polygenic prediction of height, BMI, and intelligence within families

Revisiting the limited utility of embryo screening

Nucleus Genomics Scientific Team

Abstract

In 2019, Karavani et al. (1) evaluated polygenic embryo screening using the flagship scores of the time and concluded that screening had limited utility. Nucleus Vitruvian trait models revisit that verdict. Built with SBayesRC, a Bayesian framework that analyzes ~7 million variants jointly with functional genomic annotations, they reach within-family accuracies of 43% of height variance, 18% of BMI variance, and 19% of intelligence variance. By the formula from Karavani et al., the expected spread between the highest- and lowest-scoring of five embryos using Nucleus Vitruvian trait models is ~7.5 cm (3.0 in) in height, ~10.8 IQ points, and ~5.2 kg/m² in BMI. The corresponding spreads for the 2019 scores, on the same within-family scale, were ~5.2 cm, ~5.4 IQ points, and ~3.7 kg/m². Nucleus models additionally surpass flagship academic scores trained on sample sizes up to nearly four times larger.

Nucleus Vitruvian models

The base genetic associations used in Nucleus PGS are derived from Genome-Wide Association Studies (GWAS), which test millions of common genetic variants for association with trait status. For all biobanks with available individual-level data, Nucleus performs GWAS using REGENIE (2), a machine-learning whole-genome regression framework with ridge regression and leave-one-chromosome-out (LOCO) predictions to account for polygenic background and control for population structure and relatedness. GWAS associations across multiple biobanks, including FinnGen (3), the Million Veteran Program (4), and UK Biobank (5), are meta-analyzed to improve precision of variant effect size estimates.


Nucleus trait models are then constructed from these associations with SBayesRC (6), a full-Bayesian framework that jointly analyzes ~7 million imputed common variants with 96 functional genomic annotations, rather than the ~1 million HapMap3 subset used by most published scores. In benchmarking, SBayesRC improves prediction accuracy by 14% over SBayesR and outperforms LDpred-funct, MegaPRS, and PRS-CSx, with the largest relative gains in cross-ancestry application (6). Genome coverage and functional priors more than compensate for the smaller training set: Nucleus height and BMI models are trained using association data from ~1.4 million individuals, less than a third of the training data behind the field’s current best academic scores (7,8), yet exceed their within-family performance. The Nucleus intelligence PGS was trained using association data from ~500,000 individuals.


Scores were evaluated in a cohort of ~40,000 siblings from ~20,000 sibling pairs of European ancestry in UK Biobank, with the validation individuals and their relatives above a KING kinship coefficient of 0.0442 excluded from score development. Height and weight were taken from the initial assessment-centre visit (data-fields 50 and 21002), with BMI computed as weight/height². For intelligence, performance was assessed using the first in-person administration of the UK Biobank verbal–numerical reasoning (VNR) test (data-field 20016), and adjusted for test–retest reliability following Wolfram et al. (9).

Table 1. Variance explained (R²) by source, population vs. within-family scale.

Trait

Score

Training N

Population R²

Within-family R²

Height

Karavani 2019 PGS (1)

~0.7M

0.248

0.205 (scaled)

Height

Yengo 2022 (7)

~5.4M

0.40–0.447

0.33 (measured)

Height

Nucleus

~1.4M

0.428

0.428 (measured)

BMI

Smit 2025 (8)

~5.1M

0.176

0.163 (scaled)

BMI

Nucleus

~1.4M

0.190

0.178 (measured)

IQ

Karavani 2019 PGS (Savage 2018 GWAS; tested in ASPIS) (1,10)

~0.24M

0.043 → 0.071

0.048 (scaled)

IQ

MTAG multi-trait (Lee 2018 / Allegrini 2019; tested in TEDS) (11,12)

Cognitive performance 0.26M + EA3 1.13M + highest math 0.43M + math ability 0.565M

0.11 → 0.18

0.122 (scaled)

IQ

van den Berg 2025 imputed-FI (tested in INTERVAL) (13)

~0.49M

0.058 → 0.095

0.065 (scaled)

IQ

Nucleus

~0.5M

0.238

0.191 (measured)

“Measured” = directly estimated within-family (sibling or parental-genotype-controlled) value; “scaled” = population R² × the square of the direct-to-population effect ratio reported by Okbay et al. (14) for actual PGIs (height 0.910² = 0.83, BMI 0.962² = 0.93, cognitive performance 0.824² = 0.68), used here to translate published population-scale PGSs to the within-family scale. Within-family R² denotes δ², the squared standardized direct-effect coefficient in population SD units. IQ population and within-family R² values are reported on the latent cognitive-ability scale after adjustment for test–retest reliability. No published reliabilities were found for the TEDS or ASPIS composites, so the UK Biobank VNR test–retest reliability of 0.607 (9,15) is applied to all academic scores. Because these composites each aggregate four separate tests (1,16), whereas the VNR is a single 13-item test (15), this convention likely errs toward overstating the academic scores’ accuracy. For INTERVAL, the validated phenotype is the average of two administrations of that same test taken 12 months apart (between-administration correlation 0.65 (13)); the reliability of a two-test average exceeds that of a single test (≈0.79), so applying 0.607 here likewise errs toward overstating the score’s accuracy. Parental-genotype-controlled regressions of the van den Berg score in the childhood ALSPAC cohort give within-family R² ≈ 0.026–0.029 (13), below its Okbay-scaled value here (0.065), consistent with stronger within-family attenuation in childhood cohorts. The adult Okbay ratios are applied uniformly for comparability.

For intelligence, the estimated within-family SNP heritability provides useful context for interpreting score performance. Tan et al. estimate the within-family SNP heritability of measured cognitive performance at ≈0.188, corresponding to ≈0.31 on the latent scale after adjustment for a test–retest reliability of 0.607 (9,17). The best comparable academic score captures roughly 39% of the common genetic variance in latent cognitive ability, whereas Nucleus’ within-family R² of 0.19 captures roughly 62%.


The expected spread from embryo screening can be computed directly from within-family R by multiplying the formula for expected gain from Karavani et al. by two. The expected difference between the highest- and lowest-scoring of n embryos is

E[spread] = √2 · δ · σ · E[Z(n)]

where δ is the standardized direct effect of the score (the within-family regression coefficient of phenotype on score, both in population SD units), σ is the trait’s standard deviation, and E[Z(n)] is the expected value of the best of n random draws from a standard normal distribution (1,18). The expected gain from selecting the top-scoring embryo over the average embryo is exactly half the expected spread.

Table 2. Expected spread between the highest- and lowest-scoring of n embryos.

Trait

Score (within-family R²)

δ

n = 5

n = 10

IQ

Karavani 2019 PGS (0.048)

0.22

5.4 pts

7.2 pts

IQ

MTAG multi-trait (0.122)

0.35

8.6 pts

11.4 pts

IQ

van den Berg 2025 (0.065)

0.25

6.3 pts

8.3 pts

IQ

Nucleus (0.191)

0.44

10.8 pts

14.3 pts

Height

Karavani 2019 PGS (0.205)

0.45

5.2 cm

6.9 cm

Height

Yengo 2022 (0.33)

0.57

6.6 cm

8.7 cm

Height

Nucleus (0.428)

0.65

7.5 cm

10.0 cm

BMI

Selzam 2019 PGS (0.090)

0.30

3.7 kg/m²

4.9 kg/m²

BMI

Smit 2025 (0.163)

0.40

5.0 kg/m²

6.6 kg/m²

BMI

Nucleus (0.178)

0.42

5.2 kg/m²

6.9 kg/m²

Spreads are computed with the exact formula above from each score’s within-family R² (in parentheses; σ: IQ = 15; height = 7.0 cm; BMI = 7.5 kg/m²; E[Z(5)] = 1.163, E[Z(10)] = 1.539). The expected gain from selecting the top-scoring embryo is exactly half the expected spread. Karavani et al.’s published projections (≈2.5 cm / ≈2.5 IQ points of expected gain at n = 5) used population-scale accuracy; their expected spreads have been calculated using the within-family scale. They did not analyze BMI, so the 2019 BMI row uses Selzam et al.’s directly measured within-family β for the Yengo 2018 PGS.

Discussion

In 2019, Karavani et al. evaluated polygenic embryo screening using theory, simulations of offspring from actual and randomly matched couples, and data from 28 large nuclear families (1). Their height PGS explained 24.8% of phenotypic variance (r ≈ 0.50; GWAS N ≈ 700,000 (19)) and their cognitive-ability PGS explained 4.3% (r ≈ 0.21; GWAS N ≈ 235,000 (10)). Karavani et al. derived a formula for the expected gain of selecting the highest-scoring of n embryos, which together with their simulations projected mean gains of ≈2.5 cm and ≈2.5 IQ points from selecting the highest-scoring of five embryos compared to the average embryo. This translates to expected spreads of ≈5 cm and ≈5 IQ points between the highest- and lowest-scoring of five embryos. The formula additionally establishes that trait gains and spreads are directly proportional to score r. Karavani et al. then concluded that screening human embryos for polygenic traits has limited utility.


Karavani et al.’s 2019 projections used the best PGS available at the time. Recalculating with the same within-family scaling on today’s within-family R² framework, updated scores now yield larger expected spreads. PGS accuracy is conventionally reported between families, where the score captures not only direct genetic effects, but also indirect genetic effects (genetic nurture, assortative mating, and population stratification), which are predictive of environmental variance. These indirect effects do not predict trait differences within a family where siblings share an environment. Okbay et al. directly estimated the ratio of direct to population effects for actual PGIs at 0.910 for height, 0.962 for BMI, and 0.824 for cognitive performance (14). Squaring these ratios gives 83%, 93%, and 68%, which we use as trait-level attenuation ratios to translate population-scale PGS R² to within-family R². Applying this scaling to the PGS from Karavani et al., their 24.8% height R² and 4.3% IQ R² (7.1% on the latent scale after reliability correction) correspond to within-family R² of 21% and 4.8%. These result in an expected height spread between the highest- and lowest-scoring of five embryos of ≈5.2 cm, and an expected IQ spread of ≈5.4 points. Karavani et al. did not analyze BMI; the 2019 benchmark is Selzam et al.’s within-family validation of the Yengo 2018 score (β = 0.30 in dizygotic twin pairs) (16,19), corresponding to an expected spread of ≈3.7 kg/m² among five embryos. Sibling and parental-genotype analyses of other published scores confirm substantial within-family attenuation for cognitive scores and milder attenuation for height and BMI (16,20–22).


The linear relationship between selection gain and δ means that within-family accuracy is the key predictive quantity for polygenic embryo screening. Consequently, the expected IQ spread from academia’s best score moved from ≈5.4 points in 2019 to ≈8.6 points in 2026, while height moved from ≈5.2 to ≈6.6 cm. Nucleus models yield expected spreads of ≈10.8 IQ points and ≈7.5 cm of height, twice the IQ spread, and roughly 1.5 times the height spread, projected by Karavani et al. in 2019.

Conclusion

Nucleus Vitruvian models achieve substantially higher within-family prediction than the best comparable academic scores. Across five embryos, these models deliver an expected spread of ≈10.8 IQ points, ≈7.5 cm in height, and ≈5.2 kg/m² in BMI, versus ≈8.6 points, ≈6.6 cm, and ≈5.0 kg/m² for the academic scores. All accuracies are estimated in cohorts of predominantly European ancestry, under the published reliability and within-family attenuation conventions described above.

References

  1. Karavani E, Zuk O, Zeevi D, et al. Screening human embryos for polygenic traits has limited utility. Cell. 2019;179(6):1424–1435. doi.org/10.1016/j.cell.2019.10.033

  2. Mbatchou J, Barnard L, Backman J, et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat Genet. 2021;53(7):1097–1103. doi.org/10.1038/s41588-021-00870-7

  3. Kurki MI, Karjalainen J, Palta P, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613(7944):508–518. doi.org/10.1038/s41586-022-05473-8

  4. Gaziano JM, Concato J, Brophy M, et al. Million Veteran Program: a mega-biobank to study genetic influences on health and disease. J Clin Epidemiol. 2016;70:214–223. doi.org/10.1016/j.jclinepi.2015.09.016

  5. Bycroft C, Freeman C, Petkova D, et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562(7726):203–209. doi.org/10.1038/s41586-018-0579-z

  6. Zheng Z, Liu S, Sidorenko J, et al. Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries. Nat Genet. 2024;56(5):767–777. doi.org/10.1038/s41588-024-01704-y

  7. Yengo L, Vedantam S, Marouli E, et al. A saturated map of common genetic variants associated with human height. Nature. 2022;610(7933):704–712. doi.org/10.1038/s41586-022-05275-y

  8. Smit RAJ, Wade KH, Hui Q, et al. Polygenic prediction of body mass index and obesity through the life course and across ancestries. Nat Med. 2025. doi.org/10.1038/s41591-025-03827-z

  9. Wolfram T, Moore S, Li JH, et al. Interpreting polygenic prediction of cognitive ability: evidence for direct, reliable, and portable genetic effects. Intell Cogn Abil. 2026;2(1):1–19. icajournal.scholasticahq.com/article/158459

  10. Savage JE, Jansen PR, Stringer S, et al. Genome-wide association meta-analysis in 269,867 individuals identifies new genetic and functional links to intelligence. Nat Genet. 2018;50(7):912–919. doi.org/10.1038/s41588-018-0152-6

  11. Lee JJ, Wedow R, Okbay A, et al. Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Nat Genet. 2018;50(8):1112–1121. doi.org/10.1038/s41588-018-0147-3

  12. Allegrini AG, Selzam S, Rimfeld K, et al. Genomic prediction of cognitive traits in childhood and adolescence. Mol Psychiatry. 2019;24(6):819–827. doi.org/10.1038/s41380-019-0394-4

  13. van den Berg DM, Huang W, Malawsky DS, et al. Imputation of fluid intelligence scores reduces ascertainment bias and increases power for analyses of common and rare variants. medRxiv. 2025. doi.org/10.1101/2025.06.18.25329418

  14. Okbay A, Wu Y, Wang N, et al. Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nat Genet. 2022;54(4):437–449. doi.org/10.1038/s41588-022-01016-z

  15. Fawns-Ritchie C, Deary IJ. Reliability and validity of the UK Biobank cognitive tests. PLoS ONE. 2020;15(4):e0231627. doi.org/10.1371/journal.pone.0231627

  16. Selzam S, Ritchie SJ, Pingault J-B, et al. Comparing within- and between-family polygenic score prediction. Am J Hum Genet. 2019;105(2):351–363. doi.org/10.1016/j.ajhg.2019.06.006

  17. Tan T, Jayashankar H, Guan J, et al. Family-GWAS reveals effects of environment and mating on genetic associations. medRxiv. 2024 (v3, 2026). doi.org/10.1101/2024.10.01.24314703

  18. Zuk O. EmbryoSelectionCalculator: code for expected gains in embryo selection, accompanying ref. 1. GitHub. github.com/orzuk/EmbryoSelectionCalculator

  19. Yengo L, Sidorenko J, Kemper KE, et al. Meta-analysis of genome-wide association studies for height and body mass index in ~700,000 individuals of European ancestry. Hum Mol Genet. 2018;27(20):3641–3649. doi.org/10.1093/hmg/ddy271

  20. Howe LJ, Nivard MG, Morris TT, et al. Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. Nat Genet. 2022;54(5):581–592. doi.org/10.1038/s41588-022-01062-7

  21. Young AS, Nehzati SM, Lee C, et al. Mendelian imputation of parental genotypes improves estimates of direct genetic effects. Nat Genet. 2022;54(6):897–905. doi.org/10.1038/s41588-022-01085-0

  22. Lin Y, Procopio F, Keser E, et al. Polygenic score prediction within and between sibling pairs for intelligence, cognitive abilities, and educational traits from childhood to early adulthood. Intell Cogn Abil. 2025. icajournal.scholasticahq.com/article/140654