GeneScope
Reads a consumer DNA export against ClinVar and PharmGKB, then four agents interpret it
- Role
- Solo, directing agents
- Status
- Paused
- Source
- Public on GitHub
- Stack
- TypeScriptNode.jsNext.jsClinVarPharmGKBClaude Code skills
In numbers
308K
clinically significant ClinVar variants kept from the full download
4
Claude agents interpreting in parallel, plus an orchestrator
3
consumer DNA formats detected automatically
The problem
A raw DNA file from a consumer test is about 600,000 genotype rows with no interpretation attached. The public evidence to read it exists, in ClinVar for disease variants and in PharmGKB and the CPIC guidelines for drug-gene interactions, but it is spread across databases nobody queries by hand. My degree is in pharmacology, and the drug-metabolism genes (CYP2D6, CYP2C19, DPYD, TPMT and others) are the part I studied, with published guidelines behind them.
The approach
Deterministic code does the matching and agents do the reading. A TypeScript pipeline parses the file and matches every genotype against the databases first; only then do agents research what the matches mean for this person, and an orchestrator merges what they find.
How it works
Matching before interpretation
Setup downloads ClinVar, about 415 MB, and filters it to about 308,000 clinically significant variants; PharmGKB's clinical annotations, with evidence levels 1A to 2B, ship in the repo. The genome becomes queryable JSON keyed by rsID on GRCh37. Exports from MyHeritage, 23andMe and AncestryDNA are detected and loaded natively.
An intake like a clinic
Before any analysis the tool asks what a doctor would: age, sex, ancestry, medications, family history, lifestyle and the person's own concerns. The pharmacogenomics agent checks its findings against the medications the person takes.
Four specialists at once
Four Claude agents research in parallel: pharmacogenomics, verified against CPIC and dbSNP; methylation and neurotransmitter genes, read as pathways instead of single SNPs; disease risk, with ClinVar findings checked again and set against family history; and lifestyle and nutrition. An orchestrator writes the cross-domain synthesis, and a follow-up loop asks targeted questions about the person's own results.
Local by design
The genome file, the findings and the reports stay in a folder on the machine, and nothing is uploaded. Each person gets an isolated profile folder, so two people's results never mix. A static V1 report generator stays beside V2 as a baseline.
What I chose, and what lost
Chose
An agent team that researches each finding
Over
V1's static database of hardcoded descriptions
A template prints the same paragraph for a variant whatever the person's medications or history are. The agents check current sources against the intake.
Outcome
The command-line flow through Claude Code is the working interface. A Next.js report viewer exists, built for V1's report layout and not yet updated for V2's five-file output, and the README says so. The tool is for information and education; its disclaimer rules out clinical diagnosis and sends medical decisions to a genetic counsellor or a doctor. It is paused, and personal results stay off this site.