How to Choose Sequencing Depth
Why the small number after the “x” quietly decides what a genetic test can and cannot find.
Authors: Sameer Malik, Sana
Most of the times when we connect with a lab to get the genetic sequencing done, they often ask us about the sequencing depth. Sometimes the answer is not simple, and as a geneticist, researcher, or scientist, if you are ordering the genetic test, not knowing this allows the lab to choose a default that might not work in your case. So here is the guide to make that process easy. This article explains what depth is, how it differs from the closely related idea of coverage, why it changes what a test can see, and how to decide how much of it a given case actually needs. No matter what sequencing depth you choose, you can analyze it with Vgen23.
Every sequencing report has this sequencing depth as a number that sits in the methods, written as something like 30x, 100x, or 250x. This number describes how carefully the DNA was read, and it sets a ceiling on what the test could find. A variant that explains a patient’s illness can be present in the DNA and still go unreported because that part of the DNA was not read enough times to be sure of the result.
Sequencing depth is one of the early decisions in genetic testing, made long before anyone looks at a single result. Set it too low and real findings hide in the noise. Set it higher than the question needs, and you pay for confidence you already had.
What “Depth” Actually Means
Let’s start with what the sequencer does. It does not read your DNA once, smoothly, from beginning to end. It breaks the sample into millions of short fragments, reads each one, and then the software places those reads back in their correct positions on a standard human DNA map called the reference genome. At any single position, several reads usually pile up on top of one another. Sequencing depth is simply how many reads land on a given position, averaged across the region that was tested.
So 30x means that, on average, every position was read about thirty times. 100x means about a hundred. The “x” is just shorthand for “times read.”
Why read the same letter thirty times instead of once? Because a single read is one witness, and witnesses make mistakes. Even a high-quality sequencer misreads roughly one letter in a thousand. If you read a position once and see a change, you cannot tell whether you have found a real variant or simply caught the machine in a typo. Read that same position thirty times, and the picture becomes clearer. A genuine change shows up in read after read. A random error appears once and gets outvoted by the rest. In short, depth is what allows a test to be confident about each letter it reports.
Depth Is Not The Same As Coverage
The words ‘depth’ and ‘coverage’ are often used as if they mean the same thing. This causes a lot of confusion, because they describe two different ideas.
TWO NUMBERS, TWO QUESTIONS Depth (or read depth) answers, "How many times was each position read?” It is the per-position number, reported as an average, like 100x. More depth means more confidence in each letter. Breadth of coverage answers, "How much of the target was read at all?” It is usually a percentage, such as 98% of the exome covered. More breadth means fewer gaps. The phrase “30x coverage” blurs the two together. It is really a statement about depth. When a report gives coverage as a percentage, it is talking about breadth. A good test needs both of these: enough depth to trust each call and enough breadth that the regions you care about were actually read. |
You can have one without the other. A test can hit a high average depth and still leave specific genes in the dark, because depth is an average, and averages hide their worst cases. We will come back to this, because it is where many real-world misses come from.
Why Depth Decides What You Can See
Confidence in each letter is the first reason depth matters. The second is subtler and more important: depth decides whether you can detect a variant that is present in only some of the DNA.
Most of the genome comes in two copies, one from each parent. When a variant sits on only one of those copies, it is called "heterozygous," and roughly half the reads at that position should carry it. That “roughly” matters here. If you only have a few reads at a position, chance alone can give you all reference reads and none that carry the variant, the same way three coin flips can easily come up all heads even with a fair coin. The variant is present. The test simply did not read enough times to catch it. Adding depth sharply reduces the chance of missing a true heterozygous variant.
The problem grows harder when the variant is rarer still within the sample. In cancer, a tumour mutation may be present in only a small fraction of the cells in a biopsy, so only a small fraction of reads carry it. In mosaicism, where a variant appears after conception and ends up in only some of a person’s cells, the fraction can be lower again. To see a change that appears in, say, five percent of reads, and to be sure it is real and not noise, you need many reads at that position. This single fact explains most of the gap between a 30x germline genome and a tumour panel read at 500x or deeper. The smaller the fraction of cells you need to detect, the deeper you have to read.
The Problem with Averages
Here is the trap inside that headline number. Depth is reported as an average, and an average says nothing about its worst neighborhoods. A sample at 100x average might read some regions three hundred times and others barely five, with a few stretches at zero. Some parts of the genome are simply hard to read: regions rich in the letters G and C, long repeats, and certain clinically important ones tend to come out thin no matter how high the average climbs.
This is why an average on its own can be misleading. A report can claim a healthy average and still have blind spots inside genes that matter for the patient. The more useful number is the fraction of the target read above a minimum level, written as something like “percent of targeted bases covered at 20x or higher.” A well-run clinical exome aims to push almost all of its targets above that minimum, not just to hit a big average. When you read a report, that breadth-at-a-threshold number tells you more about what could have been missed than the single headline x ever will.
WHAT TO ASK ABOUT DEPTH ON A REPORT Average depth alone is not enough. Two questions reveal more. 1. What fraction of the target was covered at the clinical minimum, for example, at 20x or higher? This reveals blind spots that an average can hide. 2. Were the specific genes relevant to this patient adequately covered? A strong genome-wide average can still sit on top of one poorly read gene of interest. |
Reading the Numbers: How Much Depth for What
Different jobs need different depths, and the typical figures below follow straight from that logic. The point is not to memorize numbers but to see why they change. In every row, the depth number is really a function of the smallest fraction of reads the test needs to see a variant in. Germline variants sit near fifty percent, so 30x is enough. Tumor and cell-free DNA variants can sit below one percent, so depth climbs by an order of magnitude or more.
| Test or application | Typical depth | Why |
|---|---|---|
| Germline whole-genome sequencing | about 30x | Looking for inherited variants present in half or all of the reads. Around 30x reliably detects single-letter changes and small insertions and deletions across the genome. |
| Germline whole-exome sequencing | about 100x average | The step that pulls out the coding regions is uneven, so a higher average is needed to lift most targets above a usable minimum. |
| Targeted gene panels | 200x to 1,000x | The region is small, so deep reading is affordable, and the extra depth allows confident calls and detection of lower-level variants. |
| Tumor and somatic panels | 250x to 1,000x or more | Tumor samples are mixtures. Mutations may sit in a small fraction of cells, and damaged biopsy tissue wastes reads, so depth has to increase. |
| Cell-free DNA and liquid biopsy | 1,000x to many thousands | Variant fractions can fall below one percent. These are visible only with very deep reading, often paired with molecular tagging to filter out errors. |
| Rapid neonatal whole-genome sequencing | about 30x to 40x | Depth stays near the germline standard, but the process is built for speed when a critically ill newborn cannot wait weeks. |
Typical figures for orientation. Exact targets vary by laboratory, platform, and assay design.
How to Choose: A Working Framework
Depth is best treated as an answer to a question, not a default copied from the last test. A few steps get you to the right number.
1. Start from the clinical question, not the machine setting. The clinical question is really three questions bundled together: what condition or category you suspect, what kinds of DNA changes typically cause it (single-letter changes, small indels, copy number changes, repeat expansions, structural rearrangements, or mosaic changes), and the lowest fraction of cells or reads in which you would need to catch the variant. Depth falls out of these three. When symptoms do not narrow to a specific disease, that is still an answer: it usually points to broad testing at a standard depth, with the understanding that you are trading specificity for breadth.
2. Match depth to the lowest allele fraction you need to detect. A heterozygous germline variant shows up in about half the reads, so 30x is comfortable. A somatic variant at five percent needs hundreds of reads. A low-level mosaic needs more again.
3. Match the method to the variant type. Once step 1 has named the likely variant types, check whether depth is even the right lever. Single-letter changes and small indels respond well to more depth. Copy number changes, repeat expansions, and large structural rearrangements depend far more on even coverage, read length, and capture design. For those, a different method beats a bigger number.
4. Factor in sample quality. Degraded DNA, tumor tissue preserved in paraffin, saliva, and low-input samples all lose usable reads. Allow extra depth so the data lands where you intended.
5. Look past the average to how evenly the DNA was read. Ask for the percent of targets above a clinical minimum and for coverage of the specific genes that matter, not just the headline figure.
6. Respect the cost curve. Depth is not free, and its returns flatten beyond the point your question requires. Extra reads mostly reconfirm what you already knew.
Where More Depth Stops Helping
It is tempting to treat depth as a dial you can always turn up for a better result. It is not. The returns flatten, and some problems do not yield to depth at all.
For a clean germline genome, moving from 30x to 60x adds a little confidence at the edges and rarely changes the diagnosis. Pushing the same case to 300x mostly spends money to re-read letters you were already sure about. Depth buys the most when the thing you are chasing is faint and very little once that signal is already clear.
Depth also cannot rescue what the method never captured. It does not fill in a region the capture step failed to pull out. It does not stretch a short read across a long repeat that short reads cannot span. It does not undo the damage in a degraded sample. Those are limits of chemistry, capture design, and read length, and the fix is a different approach, not a bigger number. This is exactly where long-read sequencing is useful, because it can resolve repeats and structural changes that short reads miss no matter how deep they go.
Depth Is a Question, Not a Setting
Strip away the jargon, and the choice is simple to state. Sequencing depth is how many times you read each position, and that number quietly sets the limit on what a test can find. Read deeply enough and faint signals come into view: a heterozygous variant that chance might have hidden, a tumour change present in only five percent of cells, or a low-level mosaic. Read too shallowly and they stay invisible, lost in the ordinary noise of the machine.
The right depth is not a fixed value to copy. It falls out of what you are trying to find and in whom. A clean inherited-disease genome is well served by around 30x. A tumor with a small subclone needs far more. Cell-free DNA needs more still. And whatever the number, the average alone never tells the whole story, because a blind spot inside the right gene can quietly undo a healthy headline figure.
A genetic test can only answer what was read deeply enough to see. Choosing depth with the question in mind and knowing how to read the depth on a report when one lands on your desk is part of making sure the answer already written in the DNA actually reaches the patient in front of you.
References
(1) Sims D, Sudbery I, Ilott NE, Heger A, Ponting CP. Sequencing depth and coverage: key considerations in genomic analyses. Nat Rev Genet. 2014;15(2):121-132. doi:10.1038/nrg3642
(2) Goodwin S, McPherson JD, McCombie WR. Coming of age: ten years of next-generation sequencing technologies. Nat Rev Genet. 2016;17(6):333-351. doi:10.1038/nrg.2016.49
(3) Meienberg J, Bruggmann R, Oexle K, Matyas G. Clinical sequencing: is WGS the better WES? Hum Genet. 2016;135(3):359-362. doi:10.1007/s00439-015-1631-9
(4) Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: ACMG/AMP joint consensus recommendation. Genet Med. 2015;17(5):405-424. doi:10.1038/gim.2015.30
(5) Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434-443. doi:10.1038/s41586-020-2308-7
(6) Newman AM, Bratman SV, To J, et al. An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage. Nat Med. 2014;20(5):548-554. doi:10.1038/nm.3519