Case study Family project
Genealogy
Applied AI · Graph model · Language processing
Rebuilding the family tree, from genealogy sites and handwritten archives spanning the 16th century to today. AI helps me trace ancestries and transcribe documents few people can still read with the naked eye.
My grandfather had scribbled a rough start: about 80 people, his parents, his siblings, the line down to us.
The urge to digitize it, and go further. AI could help scan genealogy sites (Geneanet, FamilySearch) in my place. But I want to check what it finds, so I still need the archives themselves. Once the search legwork is done, the birth, marriage, and death records are easy to track down: I know where to look, what, who, and roughly when.
These records themselves get harder to read over time: paper degradation, cursive script, sometimes Latin. AI helps here too, transcribing and checking. Any fact found gets kept, even outside direct lineage: an occupation, a place, it can still come in handy.
The engineer's reflex stays the same: from data source to visualization, through deployment, with quality checks at every step. That's how 80 people became over 1,200 today.
Past the tedious slog through genealogy sites, this is where AI has helped me the most. I feed it the archive images, and it gives me a line-by-line transcription of each handwritten record. It's on me to confirm or correct that transcription. The original spelling is always kept as written, as an image, and the proposed reading makes each individual card in the graph easier to check.
Le lundy treizième jour de janvier l'an mil six cens furent épousez François Gouaud et Gillette Gérard à Campbon
type: mariage date: 13.01.1600 location: Campbon persons: - francois-gouaud - gillette-gerard
An uncertain reading is flagged as such, never shown with the same confidence as a sure one.
A family "tree" always has one privileged root. Taken strictly, it leaves out uncles and aunts, cousins, nephews, and so on. That's not enough to represent what I want: the storage is a graph that can grow sideways.
The result: sibling groups and their children find their place, with no root forced on the graph. The "tree" you see on screen is just one walk through the graph from a chosen focus person, not the actual shape of the data.
Honestly, past a certain depth of ancestry, those sibling groups stop being interesting. Only the direct line stays in the graph.
A birth dated differently by two sources? Both dates stay on file, each with its own source and confidence level. A record read straight from the archives outweighs a crowdsourced profile on a genealogy site.
A barely legible date keeps its raw spelling. An inferred date is marked as medium confidence, with a note explaining the inference.
A remarriage with no descendants linked to the family doesn't get its own card, even with a genuine record in hand. The graph only carries what actually connects to someone.
The site stays private: the cards hold the civil status of living people who never consented to publication. No link to click here, just how it works and a few screenshots.
If the topic interests you, the architecture or the transcription approach, I'm happy to talk about it.
Get in touch →References
-
[01]
Geneanet
Crowdsourced family trees and partner-digitized record collections. The starting point for spotting a lead before checking it against the primary record.
-
[02]
FamilySearch
Free, massive collection run by The Church of Jesus Christ of Latter-day Saints. Especially strong on older registers and branches outside France.
-
[03]
France Archives
The national portal aggregating digitized holdings from France's departmental archives: birth, marriage and death records, parish registers. The primary source for any French genealogist, town by town.