Discover the Hidden Dangers Lurking in Your FIGG SNP Profile

Discover the Hidden Dangers Lurking in Your FIGG SNP Profile

Forensic investigative genetic genealogy depends on the assumption that a SNP profile uploaded to a genealogy database will produce a reliable match list. In practice, a genotype file can look usable while hiding technical problems that directly affect an investigation. False-negative segments can reduce, mis-rank, or completely remove a significant relative from the match list, causing investigators to miss the very lead needed to advance a case. False-positive segments can make a profile appear artificially related to someone it is not, sending investigators down a rabbit hole chasing a match who has no biological connection to the sample. These problems may arise from a litany of causes such as: low coverage, reduced callable-SNP density, allelic dropout, high genotype call error rates, cytosine deamination, contamination, wet-lab artifacts, sequencing artifacts, batch effects, and inappropriate genotype filtering. Depending on how they affect genotype calls and shared-segment detection, these issues can cause sensitivity loss, excess “matchiness,” or both.

The consequences are not limited to obvious upload failures. On GEDmatch, the most visible failure mode is a kit being marked “matchy,” but bad files may also trigger heterozygosity-based filters. However, successful upload does not guarantee a trustworthy match list. A kit can avoid the explicit “matchy” designation while still producing excess false-positive segments that distort the apparent relationship landscape. Likewise, an apparently sparse match list may reflect technical sensitivity loss rather than the absence of useful relatives.

This presentation will describe two complementary tools Astrea Forensics has built to evaluate FIGG genotype quality and performance. SNPShake analyzes an evidentiary genotype file for potential problems before or after database upload, examining marker-set compatibility, missingness, heterozygosity patterns, damage-associated signatures, contamination signals, and matchiness-related metrics. By contrast, astrea-sensitivity-test is designed to benchmark a bioinformatics pipeline when both known-truth data, such as high-quality sequencing or microarray genotypes, and casework-like data are available. It compares the genotype file produced from the casework-like input against the known truth to measure sensitivity, genotype error, allelic dropout, false-positive and false-negative segment behavior, and database-relevant SNP overlap. Together, these tools help distinguish biological reality from technical artifact and support decisions about whether a case should proceed, be reanalyzed, or be regenerated from sequencing data.

We will also present Astrea’s optimized whole-genome sequencing FIGG pipeline, WOPR, built to generate high-quality genealogy-compatible genotype files from challenging forensic samples. WOPR is a reproducible, auditable, containerized workflow with genotype filtering designed to reduce matchy artifacts and heterozygosity-filter failures, producing FIGG-optimized genotype files, reports, sex inference, mtDNA haplogroup calls, and detailed QC metrics. Validation under ISO 17025 and real-world GEDmatch testing demonstrated interpretable results at 0.1x sequencing coverage, with minimum deliverable files containing more than 460,000 SNPs, and successful performance on difficult sample types including bone and rootless hair. Attendees will leave with practical guidance for recognizing hidden SNP-profile hazards, understanding how those hazards affect an investigation, and determining when optimized WGS processing can rescue or improve a FIGG case.

Forensic investigative genetic genealogy depends on the assumption that a SNP profile uploaded to a genealogy database will produce a reliable match list. In practice, a genotype file can look usable while hiding technical problems that directly affect an investigation. False-negative segments can reduce, mis-rank, or completely remove a significant relative from the match list, causing investigators to miss the very lead needed to advance a case. False-positive segments can make a profile appear artificially related to someone it is not, sending investigators down a rabbit hole chasing a match who has no biological connection to the sample. These problems may arise from a litany of causes such as: low coverage, reduced callable-SNP density, allelic dropout, high genotype call error rates, cytosine deamination, contamination, wet-lab artifacts, sequencing artifacts, batch effects, and inappropriate genotype filtering. Depending on how they affect genotype calls and shared-segment detection, these issues can cause sensitivity loss, excess “matchiness,” or both.

The consequences are not limited to obvious upload failures. On GEDmatch, the most visible failure mode is a kit being marked “matchy,” but bad files may also trigger heterozygosity-based filters. However, successful upload does not guarantee a trustworthy match list. A kit can avoid the explicit “matchy” designation while still producing excess false-positive segments that distort the apparent relationship landscape. Likewise, an apparently sparse match list may reflect technical sensitivity loss rather than the absence of useful relatives.

This presentation will describe two complementary tools Astrea Forensics has built to evaluate FIGG genotype quality and performance. SNPShake analyzes an evidentiary genotype file for potential problems before or after database upload, examining marker-set compatibility, missingness, heterozygosity patterns, damage-associated signatures, contamination signals, and matchiness-related metrics. By contrast, astrea-sensitivity-test is designed to benchmark a bioinformatics pipeline when both known-truth data, such as high-quality sequencing or microarray genotypes, and casework-like data are available. It compares the genotype file produced from the casework-like input against the known truth to measure sensitivity, genotype error, allelic dropout, false-positive and false-negative segment behavior, and database-relevant SNP overlap. Together, these tools help distinguish biological reality from technical artifact and support decisions about whether a case should proceed, be reanalyzed, or be regenerated from sequencing data.

We will also present Astrea’s optimized whole-genome sequencing FIGG pipeline, WOPR, built to generate high-quality genealogy-compatible genotype files from challenging forensic samples. WOPR is a reproducible, auditable, containerized workflow with genotype filtering designed to reduce matchy artifacts and heterozygosity-filter failures, producing FIGG-optimized genotype files, reports, sex inference, mtDNA haplogroup calls, and detailed QC metrics. Validation under ISO 17025 and real-world GEDmatch testing demonstrated interpretable results at 0.1x sequencing coverage, with minimum deliverable files containing more than 460,000 SNPs, and successful performance on difficult sample types including bone and rootless hair. Attendees will leave with practical guidance for recognizing hidden SNP-profile hazards, understanding how those hazards affect an investigation, and determining when optimized WGS processing can rescue or improve a FIGG case.

Workshop currently at capacity. A waitlist is available to join on our registration page.

Brought to you by

Worldwide Association of Women Forensic Experts

Kevin Lord

Director of Bioinformatics, Astrea Forensics

Kevin Lord is Director of Bioinformatics at Astrea Forensics, where he leads development of bioinformatics tools and workflows for forensic investigative genetic genealogy. He is the architect and lead developer of Astrea’s ISO 17025-validated WOPR pipeline, which converts challenging low-coverage and degraded sequencing data into FIGG-optimized SNP profiles.

Speaker Image

Submit Question to a speaker