Phylogeny × Data Lineage

What Phylogenetics Can Teach Data Governance

Tim Mutton Founder, Celeste IQ August 24, 2026

Long before Darwin, biologists had a problem that looks a lot like ours.

They had organisms. They wanted ancestry. Nobody had footage of evolution happening - what they had were bones, teeth, wing shapes, and the job was to work backward from those to a plausible family tree. That's phylogenetics: reconstructing lineage you can't observe directly, from whatever signals you can.

Anyone who's tried to map lineage across a real enterprise data estate will recognize this. It's the same problem.

The record you wish existed doesn't

In an ideal world, every table and field would carry its own paper trail - this value came from that source, moved through this pipeline, got renamed here. Full provenance, always current.

That world doesn't exist anywhere we've worked, and Celeste IQ's team has spent close to a decade inside utility data environments where getting this wrong has real consequences. What actually exists looks more like a fossil record: partial, missing exactly the transitional evidence you'd want most. A pipeline gets rewritten and the old logic disappears. A field gets renamed mid-migration and the catalog never catches up. A "temporary" transformation from 2019 is still load-bearing in 2026.

The lineage happened, nobody wrote it down. So the question becomes the one Cuvier and Owen and Darwin were working with: given what's left, what's the best-supported reconstruction?

A trap both fields fall into

Comparative anatomists had to learn to tell apart two kinds of similarity. A bat's wing and a human arm share the same underlying bone structure, just repurposed - homology, similarity from shared origin. A bird's wing and an insect's wing do the same job and look alike, but evolved independently - analogy, similarity from shared function with no shared history. Mix the two up and the tree falls apart.

Data has the same trap. Two fields named customer_id in different systems can look identical. Sometimes one really is derived from the other. Sometimes two teams just landed on the same naming convention without either system ever touching the other's data. Treating a matching field name as lineage evidence is the bat-wing mistake at enterprise scale.

What actually separates the two cases isn't the surface resemblance - it's what's underneath. Value distributions, transformation patterns, how a field's data behaves as it moves through a pipeline. Much harder to fake than a name.

Confidence, not certainty

The other thing phylogeny gets right: nobody presents a phylogenetic tree as settled fact. It's the best-supported hypothesis given current evidence, with a confidence level attached to each branch, revised the moment better evidence shows up.

Most lineage tooling doesn't work that way. It treats lineage as binary - mapped or not mapped. Either teams spend months on manual tagging to get certainty, or there's nothing until someone does that work.

The better approach, and the one we built Celeste IQ's lineage engine around, looks more like how biologists actually operate: infer lineage from structural signal first, attach a confidence score, and let manual review raise that confidence rather than requiring it up front. You don't sit on a complete fossil record before publishing a tree.

Why it matters

If lineage only counts once it's manually documented, you're permanently behind - real data estates change faster than any team can hand-tag them. Treat lineage the way biology treats ancestry instead: inferable from evidence, stated with confidence rather than certainty, open to revision as more comes in. What you get is usable right away, and it improves over time instead of going stale the day it's finished.

The fossil record was never going to be complete. Biology built a rigorous discipline on top of it anyway. Data governance can do the same, once it stops waiting on documentation that was never going to arrive in time.

Celeste IQ's predictive lineage engine applies this structural-signal-first approach across utility data estates - surfacing likely lineage before manual tagging catches up, with a confidence score attached to every prediction.

Explore the platform