Digital Hell All articles
Digital Dystopia

Grandma's Photo Album Is Now a Facial Recognition Dataset

Digital Hell

Somewhere in a data center, a machine is looking at a photograph of your grandfather at his high school graduation in 1957. It is learning the geometry of his face — the distance between his eyes, the angle of his jaw, the particular curve of his brow. And because faces are hereditary, it might be learning something about yours.

This is not a hypothetical. This is the current state of genealogy platforms, family photo archives, and the AI companies that have discovered something remarkable: the most intimate, historically rich, and legally underprotected facial recognition dataset in the world was assembled voluntarily by people who just wanted to know where their great-grandparents came from.

The Ancestry Gold Rush

The genealogy industry had a legitimately beautiful pitch. Upload your family photos, build your family tree, connect with distant relatives, recover the history that time and migration and tragedy tried to erase. For immigrant families, for descendants of enslaved people, for adoptees, for anyone who grew up with gaps in their family story, these platforms offered something genuinely meaningful.

Ancestry, MyHeritage, FamilySearch, 23andMe's family features — these platforms accumulated hundreds of millions of photographs. Old photos. Scanned prints. Images that had never been digitized before, uploaded by users who were doing the emotional labor of family archivism and sharing it with the world.

The photos span generations. They include faces that would never otherwise appear in any digital database — people who died before the internet existed, people who lived in rural communities that no corporate camera ever visited, people whose images exist only in shoeboxes in attics and, now, on genealogy servers.

From a facial recognition training perspective, this is extraordinary. Most facial recognition datasets are modern, often scraped from social media, heavily weighted toward certain demographics and lighting conditions and photographic styles. Genealogy archives are the opposite: diverse across time, geographically varied, spanning a century of photographic history. They are, from a machine learning standpoint, incredibly valuable.

The Consent Problem Has Several Layers

Here's where it gets ethically complicated in ways that go beyond the usual tech privacy concerns.

When you upload a photo to a genealogy platform, you are typically consenting on behalf of yourself. The terms of service you agreed to govern your data. But that photograph almost certainly contains other people — people who are dead, people who are living relatives who never agreed to anything, children who cannot legally consent, and crucially, people from generations past who had no concept of digital technology, let alone facial recognition AI, when they were alive.

The consent framework that governs most digital privacy law was not designed for this. GDPR in Europe has some provisions that complicate this, which is part of why European genealogy data practices look different from American ones. In the US, the legal landscape is significantly more permissive. There is no federal biometric privacy law. State-level protections like Illinois's BIPA are meaningful but limited in geographic scope. The FTC has taken some action in this space but nothing that comprehensively addresses the genealogy platform problem.

What this means practically is that your grandmother, photographed at her wedding in 1948, has no legal protection preventing her image from being used to train a facial recognition model. She didn't consent. She couldn't consent. And there's currently no law that says her descendants' consent is required either.

What the Companies Are Actually Doing With It

This is where the reporting gets murky, because the companies involved are not exactly forthcoming.

MyHeritage launched a feature called LiveStory that animates old photographs — making still images appear to move and breathe — using AI. The technology is impressive and deeply unsettling in equal measure. To build it, they needed to train models on exactly the kind of historical face data their platform had accumulated. The privacy implications of that training process were not extensively disclosed to users.

Ancestry's terms of service grant the company a broad license to use uploaded content. The specific language around how that content might be used in AI development has been updated multiple times, usually in ways that expand rather than restrict the company's rights.

Beyond the platforms themselves, there's the scraping problem. Academic researchers and private AI companies have scraped genealogy sites and family photo archives from across the web. Clearview AI — the facial recognition company that built a database of billions of images by scraping the internet without permission — has been the subject of multiple lawsuits and regulatory actions, but it is far from the only entity doing this kind of work. Smaller companies operate in this space with even less visibility.

The specific horror here is that genealogy photos are often labeled. Users add names, dates, locations, relationships. A scraped genealogy photo doesn't just give an AI a face — it gives an AI a face with a name and a family tree attached. That's a categorically more powerful dataset than an anonymous scraped image.

The Baby Picture Problem

Let's make this personal, because it should be.

Millions of Americans have uploaded childhood photos of themselves to genealogy platforms — either directly or because a parent or grandparent included them in a family archive. Photos from the 1980s and 1990s, when most of those people were minors. Those images are now potentially part of training datasets for facial recognition systems.

The child in that photo did not consent. The adult that child became may not even know the photo is online. And the face in that photo, combined with the face of every other family member in the surrounding images, creates a biometric fingerprint that links across generations.

Facial recognition systems trained on this data could theoretically identify people not just as adults, but trace family relationships, predict what children will look like as adults, or identify individuals through family resemblance even when a direct match isn't available. This is not science fiction. The technical capability exists.

Where This Goes From Here

The genealogy platform reckoning is coming, but it's moving slowly because the people most affected often don't realize they're affected.

Advocates in the biometric privacy space have been raising alarms about this for years. The policy apparatus moves at its usual glacial pace. The companies have little financial incentive to restrict data uses that may be generating significant revenue or competitive advantage.

What would actually help: a federal biometric privacy law with teeth. Explicit opt-in requirements for AI training use of uploaded images. Retroactive deletion rights that actually work. And a broader cultural conversation about what it means to upload not just your own face, but the faces of everyone who came before you.

You wanted to find your roots. You didn't sign up to hand the future a map of your family's faces.

But here we are.

All articles

Related Articles

Everything You Deleted Is Still Whispering Somewhere

The 4 AM Reddit Confessional: Where Exhaustion Becomes Honesty

The 4 AM Reddit Confessional: Where Exhaustion Becomes Honesty

Haunted by Your Own History: How Platforms Keep Serving You Ghosts of Connections You Already Buried

Haunted by Your Own History: How Platforms Keep Serving You Ghosts of Connections You Already Buried