Methods and limits

How we read your DNA, and what we refuse to claim.

Genetic interpretation is easy to overstate, and a polished interface can create more confidence than the evidence behind it supports. So here is the whole method in the open: where each finding comes from, how markers get combined, what we take out of the report on purpose, and exactly where the science stops. None of it is behind the paywall.

Knowledge base as of August 2026 · 1,145 curated entries across 609 markers · 325 of them cleared to reach a report

Where the findings come from

Every statement in the report traces to a public, versioned source, and every card links that source so you can read it instead of taking our word for it.

GWAS CATALOG

Effect alleles, effect sizes and p-values for common-variant associations, read from the EBI and NHGRI catalog rather than retyped from a news write-up.

CLINVAR

Reviewed classifications for cataloged variants. We carry ClinVar's review status onto the card verbatim, with its 0-4 star rating, so you can see whether a classification came from an expert panel, a single submitter, or submitters who disagree.

PHARMGKB · CPIC

Response annotations carrying their own evidence level. Only the ones about caffeine, alcohol and nicotine reach your report; everything about prescription drugs is filtered out before the report is built.

OPEN TARGETS

Gene-to-trait links used as supporting context. Never the sole basis for a finding.

PUBMED

The primary paper behind each curated statement, linked on the card so you can read the source instead of trusting our summary of it.

How a marker becomes a finding

STEP 01

Anchor the effect allele

We store which allele the research actually measured, and which direction it pushes the trait, as two separate facts. The alternate allele is simply the non-reference one and is not automatically the worse one, so a report that treats ALT as "the bad letter" gets roughly half its cards backwards.

STEP 02

Write the reading per genotype

The sentence you read is curated for your exact genotype, not generated from the numbers at display time. That keeps the prose and the data from drifting apart, which is the failure mode that produces confident text about the wrong allele.

STEP 03

Decide whether it is shown at all

Three cuts happen before you see anything. 481 of the 1,145 entries are never shown to anyone: cancer risk, chemotherapy toxicity, diagnosis-adjacent claims. 269 more are prescription-drug pharmacogenomics, stripped from every surface. 70 sit in categories the report does not cover. That leaves 325 entries across 269 markers and 112 topics, and your own file decides how many of those it actually carries.

The part everyone gets wrong: combining markers

When several markers speak to one topic, we combine them into a single lean. That combination is not a polygenic risk score, and we do not dress it up as one. The difference is worth being precise about.

A published PRSWhat Helisoma computes
PurposeEstimate risk in a defined populationOrder which findings deserve your attention first
InputsThousands to millions of variants, genome-wideThe handful of curated markers your file carries for that topic
ValidationHeld-out cohort, calibrated per ancestryNone, which is exactly why we never print a risk number
OutputA score with a population distributionThree words: favorable, typical, worth a review

Inside that combination, four rules do the real work. They exist because the naive version, adding up good and bad markers, is wrong in ways that are easy to miss and impossible to spot from the finished page.

Correlated markers count once. Markers inherited together as a block report the same underlying signal, not independent evidence. Each block collapses to a single vote. Without this rule, 11 markers in and around FTO supplied 49% of the entire body-weight lean by themselves.
Odds ratios and betas never enter one sum. An odds ratio is a multiplier; a beta is a shift measured in kg/m², mmHg, µmol/L or standard deviations. They share no scale, so each unit class is weighed on its own axis and the two are never added together.
Weights are relative, not absolute. Every signal is scored against the average effect size inside its own topic. The result is dimensionless. It tells you which way a topic leans and roughly how firmly, and it deliberately cannot be read as a quantity.
A close call is called typical. When the net lean sits within 20% of the total weight in a topic, the answer is "typical" rather than a direction we manufactured out of noise.
Nothing is a percentile. The internal ordering number carries no population distribution and is never rendered against a 50th-percentile marker. It answers "read this one first", not "you are in the Nth percentile".

153 linkage blocks are defined across the knowledge base for the first rule alone.

What we keep separate

A high-impact variant must never end up buried inside a trait score. Three kinds of genetic information behave nothing alike, so they travel through the report on separate tracks and are never summed together.

COMMON VARIANTS

Small leans measured across large groups. These are what the report is actually about, and they are stated as leans, not verdicts. One of them changes a habit at most.

CATALOGED HIGH-IMPACT VARIANTS

Most are removed outright: the 481 hidden entries are largely this material. APOE ε4 is gated instead of hidden, which means you are told the variant is present and pointed at a clinician or genetic counselor, while the verdict is forced to neutral and the effect size, direction and copy count are stripped out of the payload before the interface, the PDF or your AI ever sees them, so no risk figure can leak into a card by accident. A short list of well-characterized variants is reported normally, because leaving them out would be its own kind of dishonesty: Factor V Leiden, prothrombin G20210A, and HFE C282Y. Those carry the clinical disclaimer, an explicit note that any risk shown is relative and not your personal probability, and ClinVar's own review status, which for two of the three reads "conflicting classifications". They are a prompt to talk to a doctor and get a real test, not a substitute for one.

PRESCRIPTION-DRUG RESPONSE

Not in the report at all. We curate it, we hold the CPIC and PharmGKB evidence levels for it, and we strip all 269 of those entries out of every surface you can reach: the interactive report, the PDF, and what your AI is served. How you metabolize a prescription belongs in a conversation with whoever writes it, backed by a clinical test, not in a wellness product you bought online. What stays is substance response: caffeine, alcohol, and nicotine.

Ancestry, and why you will not see odds ratios

Most of the effect sizes behind common-variant research were measured in European-ancestry cohorts. That is a fact about the field rather than about you, and it changes how much any single number is worth.

What the skew actually does. Allele frequencies differ between populations, and the marker on your chip is often a stand-in sitting near the causal variant rather than the causal variant itself. How near it sits differs by population too. So an effect measured in one cohort can be smaller, larger, or absent in another.
What we do about it. We describe effects as small leans instead of printing odds ratios and p-values, because those figures carry a precision that does not survive the move to a different population. The report repeats this in its own Methods and limits section, and the coaching instructions tell your AI to flag the added uncertainty for non-European ancestry.
What we do not claim. There is no per-population recalibration here, and we are not going to pretend otherwise. If your ancestry is not primarily European, read every lean in the report as looser still.

What your file can and cannot deliver

Consumer arrays, exomes and whole genomes all feed the same pipeline, and each brings its own limits. These are ours.

Arrays are good at the common variants we use. On these markers a consumer chip agrees with full lab sequencing about 99% of the time. The unreliability you may have read about concerns rare disease variants, which arrays genuinely struggle with, and we do not report those as findings.
We never impute. A marker missing from your file produces no finding. We do not fill the gap statistically, because an imputation error would be invisible to you and would still read like a result.
Coverage is measured, and stated per topic. Your account shows what share of the report’s markers your file carries overall. More usefully, each topic says it directly: a chronotype section that reads 15 markers and got 3 of them prints “3/15 markers”, so you can see which conclusions rest on a fraction of their evidence instead of trusting a single global percentage.
Strand and ploidy are normalized. Everything is expressed on the GRCh38 forward strand, palindromic markers are resolved rather than guessed, and a male’s non-pseudoautosomal X is read as one copy instead of two.
We take the file as given. A genotype miscalled inside your raw file is invisible to us. We can normalize formats and strands; we cannot re-run your lab.

What the AI is allowed to do with it

The report is built to be handed to your own assistant, which is the point at which a careless product would let free-form medical interpretation back in through the side door. Four rules ship with the connector.

It has to read your actual report. The coaching instructions require the assistant to call the report tool and work from the interpretations and directions it returns, rather than answering from what it remembers about a gene.
It cannot re-derive good from bad. The assistant is told to rely on the report’s own stated direction for each finding, not to reason backwards from a raw association it half-recalls.
It gives zero interpretation on high-impact genes. For anything monogenic or ClinVar-pathogenic the instruction is to acknowledge presence and refer to a clinician: no penetrance, no risk estimate, no treatment advice.
It never sees your raw file. The connector serves the finished report and the genotypes behind it. Your raw DNA file never left your device in the first place. Once connected, that assistant processes the report under its own provider’s privacy policy, and you can disconnect at any time.

Your data

Your raw DNA file is parsed by your own browser and never uploaded. Only the small set of markers the report uses is sent to us, well under a megabyte, and we never receive or store the file itself. We do not sell personal data. You can delete your genetic data and your account yourself in settings, and the extracted markers and generated reports are deleted within 30 days. Transaction records are kept for six years under UK tax law and contain no genetic data. The full detail is in the Privacy Policy.

What this is not

It is not a diagnostic or clinical genetic test.
It is not a validated polygenic risk score, and there is no validation cohort behind it.
It is not calibrated to your ancestry, and we do not claim it is.
It does not read CNVs, structural variants, or rare variants.
It does not tell you whether you will develop a disease.
It cannot see what your file did not measure, and it says which topics that affected.

Nothing in the report should change a medication, a diagnosis, or a large health decision on its own. It is educational, it is about everyday tendencies, and the interesting questions it raises are good ones to bring to a doctor.

Questions people ask

Is the topic score a polygenic risk score?

No, and we do not present it as one. A published PRS is fitted on a discovery cohort, validated on a held-out cohort, and calibrated per ancestry, which lets it output a risk figure. Ours does none of that. It is an ordering over the handful of markers your file happens to carry, meant to answer "what should I look at first", and it outputs three words rather than a number: favorable, typical, or worth a review.

Do you report pathogenic variants, BRCA, or anything a genetic counselor would handle?

No. High-stakes clinical claims are removed from the report entirely: 481 of the 1,145 entries in our knowledge base are never shown to anyone, including cancer risk, chemotherapy toxicity, and diagnosis-adjacent findings. Where a high-impact gene appears at all, you get the fact that the variant is present plus a pointer to a clinician, with the effect size and direction stripped out of the data before it reaches the interface. If you want those answers, you want a clinical test, not a consumer report.

Do you tell me how I respond to my medications?

No. We curate prescription-drug pharmacogenomics and hold the CPIC and PharmGKB evidence levels for it, and then we strip all 269 of those entries out of the interactive report, the PDF, and anything your AI is served. Dosing is a conversation with whoever writes your prescription, backed by a clinical test. What the report does read is substance response: caffeine, alcohol, and nicotine.

How do I know how much review is behind a ClinVar classification?

The card tells you. ClinVar rates its own classifications 0 to 4 stars by how much assertion criteria stood behind them, and we print that rating with ClinVar's wording, unedited: "reviewed by expert panel" and "practice guideline" at the top, "criteria provided, single submitter" low down, "conflicting classifications" when submitters disagree, "no assertion criteria provided" at the bottom. It rates the review, not the size of the effect. Two of the three high-impact variants we do report carry conflicting classifications, and the card says so rather than rounding it up to settled.

Do you impute missing genotypes?

Never. If a marker is not in your file, it produces no finding at all. Imputation would let us claim coverage we do not have, and an imputation error is invisible to the person reading the report.

Does this work for exome or whole-genome data?

Yes, through standard VCF. The report reads the same marker set either way, so sequencing mostly buys you fuller coverage rather than different findings. Multi-gigabyte whole-genome VCFs are too large to read in a browser, so ask your provider for a sites-only or array-style export.

Do you read CNVs or structural variants?

No. Consumer arrays are not built for them and our pipeline does not call them. Everything in the report is a single-nucleotide variant read directly from your file.

Can I see the method applied to a real report before paying?

Yes. The sample report is a full report built from a sample genome, free and without an email, and it carries its own Methods and limits section.

Guides from the blog

All articles →
Helisoma Report 23andMe · AncestryDNA · MyHeritage · Nebula

Curious what your DNA says?

Your browser reads your raw DNA file right on your device, so the file itself never leaves it. Out comes a personalized report you can read yourself or hand to your AI.

SleepCaffeineNutritionTraining
Get your report →
$49 once, no subscription
free analysis first