How we read your DNA, and what we refuse to claim.
Genetic interpretation is easy to overstate, and a polished interface can create more confidence than the evidence behind it supports. So here is the whole method in the open: where each finding comes from, how markers get combined, what we take out of the report on purpose, and exactly where the science stops. None of it is behind the paywall.
Where the findings come from
Every statement in the report traces to a public, versioned source, and every card links that source so you can read it instead of taking our word for it.
Effect alleles, effect sizes and p-values for common-variant associations, read from the EBI and NHGRI catalog rather than retyped from a news write-up.
Reviewed classifications for cataloged variants. We carry ClinVar's review status onto the card verbatim, with its 0-4 star rating, so you can see whether a classification came from an expert panel, a single submitter, or submitters who disagree.
Response annotations carrying their own evidence level. Only the ones about caffeine, alcohol and nicotine reach your report; everything about prescription drugs is filtered out before the report is built.
Gene-to-trait links used as supporting context. Never the sole basis for a finding.
The primary paper behind each curated statement, linked on the card so you can read the source instead of trusting our summary of it.
How a marker becomes a finding
Anchor the effect allele
We store which allele the research actually measured, and which direction it pushes the trait, as two separate facts. The alternate allele is simply the non-reference one and is not automatically the worse one, so a report that treats ALT as "the bad letter" gets roughly half its cards backwards.
Write the reading per genotype
The sentence you read is curated for your exact genotype, not generated from the numbers at display time. That keeps the prose and the data from drifting apart, which is the failure mode that produces confident text about the wrong allele.
Decide whether it is shown at all
Three cuts happen before you see anything. 481 of the 1,145 entries are never shown to anyone: cancer risk, chemotherapy toxicity, diagnosis-adjacent claims. 269 more are prescription-drug pharmacogenomics, stripped from every surface. 70 sit in categories the report does not cover. That leaves 325 entries across 269 markers and 112 topics, and your own file decides how many of those it actually carries.
The part everyone gets wrong: combining markers
When several markers speak to one topic, we combine them into a single lean. That combination is not a polygenic risk score, and we do not dress it up as one. The difference is worth being precise about.
| A published PRS | What Helisoma computes | |
|---|---|---|
| Purpose | Estimate risk in a defined population | Order which findings deserve your attention first |
| Inputs | Thousands to millions of variants, genome-wide | The handful of curated markers your file carries for that topic |
| Validation | Held-out cohort, calibrated per ancestry | None, which is exactly why we never print a risk number |
| Output | A score with a population distribution | Three words: favorable, typical, worth a review |
Inside that combination, four rules do the real work. They exist because the naive version, adding up good and bad markers, is wrong in ways that are easy to miss and impossible to spot from the finished page.
153 linkage blocks are defined across the knowledge base for the first rule alone.
What we keep separate
A high-impact variant must never end up buried inside a trait score. Three kinds of genetic information behave nothing alike, so they travel through the report on separate tracks and are never summed together.
Small leans measured across large groups. These are what the report is actually about, and they are stated as leans, not verdicts. One of them changes a habit at most.
Most are removed outright: the 481 hidden entries are largely this material. APOE ε4 is gated instead of hidden, which means you are told the variant is present and pointed at a clinician or genetic counselor, while the verdict is forced to neutral and the effect size, direction and copy count are stripped out of the payload before the interface, the PDF or your AI ever sees them, so no risk figure can leak into a card by accident. A short list of well-characterized variants is reported normally, because leaving them out would be its own kind of dishonesty: Factor V Leiden, prothrombin G20210A, and HFE C282Y. Those carry the clinical disclaimer, an explicit note that any risk shown is relative and not your personal probability, and ClinVar's own review status, which for two of the three reads "conflicting classifications". They are a prompt to talk to a doctor and get a real test, not a substitute for one.
Not in the report at all. We curate it, we hold the CPIC and PharmGKB evidence levels for it, and we strip all 269 of those entries out of every surface you can reach: the interactive report, the PDF, and what your AI is served. How you metabolize a prescription belongs in a conversation with whoever writes it, backed by a clinical test, not in a wellness product you bought online. What stays is substance response: caffeine, alcohol, and nicotine.
Ancestry, and why you will not see odds ratios
Most of the effect sizes behind common-variant research were measured in European-ancestry cohorts. That is a fact about the field rather than about you, and it changes how much any single number is worth.
What your file can and cannot deliver
Consumer arrays, exomes and whole genomes all feed the same pipeline, and each brings its own limits. These are ours.
What the AI is allowed to do with it
The report is built to be handed to your own assistant, which is the point at which a careless product would let free-form medical interpretation back in through the side door. Four rules ship with the connector.
Your data
Your raw DNA file is parsed by your own browser and never uploaded. Only the small set of markers the report uses is sent to us, well under a megabyte, and we never receive or store the file itself. We do not sell personal data. You can delete your genetic data and your account yourself in settings, and the extracted markers and generated reports are deleted within 30 days. Transaction records are kept for six years under UK tax law and contain no genetic data. The full detail is in the Privacy Policy.
What this is not
Nothing in the report should change a medication, a diagnosis, or a large health decision on its own. It is educational, it is about everyday tendencies, and the interesting questions it raises are good ones to bring to a doctor.
Questions people ask
Is the topic score a polygenic risk score?
No, and we do not present it as one. A published PRS is fitted on a discovery cohort, validated on a held-out cohort, and calibrated per ancestry, which lets it output a risk figure. Ours does none of that. It is an ordering over the handful of markers your file happens to carry, meant to answer "what should I look at first", and it outputs three words rather than a number: favorable, typical, or worth a review.
Do you report pathogenic variants, BRCA, or anything a genetic counselor would handle?
No. High-stakes clinical claims are removed from the report entirely: 481 of the 1,145 entries in our knowledge base are never shown to anyone, including cancer risk, chemotherapy toxicity, and diagnosis-adjacent findings. Where a high-impact gene appears at all, you get the fact that the variant is present plus a pointer to a clinician, with the effect size and direction stripped out of the data before it reaches the interface. If you want those answers, you want a clinical test, not a consumer report.
Do you tell me how I respond to my medications?
No. We curate prescription-drug pharmacogenomics and hold the CPIC and PharmGKB evidence levels for it, and then we strip all 269 of those entries out of the interactive report, the PDF, and anything your AI is served. Dosing is a conversation with whoever writes your prescription, backed by a clinical test. What the report does read is substance response: caffeine, alcohol, and nicotine.
How do I know how much review is behind a ClinVar classification?
The card tells you. ClinVar rates its own classifications 0 to 4 stars by how much assertion criteria stood behind them, and we print that rating with ClinVar's wording, unedited: "reviewed by expert panel" and "practice guideline" at the top, "criteria provided, single submitter" low down, "conflicting classifications" when submitters disagree, "no assertion criteria provided" at the bottom. It rates the review, not the size of the effect. Two of the three high-impact variants we do report carry conflicting classifications, and the card says so rather than rounding it up to settled.
Do you impute missing genotypes?
Never. If a marker is not in your file, it produces no finding at all. Imputation would let us claim coverage we do not have, and an imputation error is invisible to the person reading the report.
Does this work for exome or whole-genome data?
Yes, through standard VCF. The report reads the same marker set either way, so sequencing mostly buys you fuller coverage rather than different findings. Multi-gigabyte whole-genome VCFs are too large to read in a browser, so ask your provider for a sites-only or array-style export.
Do you read CNVs or structural variants?
No. Consumer arrays are not built for them and our pipeline does not call them. Everything in the report is a single-nucleotide variant read directly from your file.
Can I see the method applied to a real report before paying?
Yes. The sample report is a full report built from a sample genome, free and without an email, and it carries its own Methods and limits section.
Guides from the blog
All articles →Curious what your DNA says?
Your browser reads your raw DNA file right on your device, so the file itself never leaves it. Out comes a personalized report you can read yourself or hand to your AI.
free analysis first