Bloodlines

What tagging given names tells us about company history (and about data modeling)

By Rob Schrauwen, former Chief Content Architect at Elsevier. Originally written September 2026 · Published here 7 October 2026

What is a given name? Two authoritative documents from the same company give opposite answers – and the reason lies in where each came from.

Although most authors now use their first name, there are authors (and journals) that continue to prefer to use initials. What would the ce:given-name, the XML tag, contain if the author name appears as "A.N. Author"?

You might expect this to be an easy question. After all, ce:given-name comes from the common element pool (the prefix ce: stands for "common element"), introduced precisely so that an element would have the same meaning across Elsevier's XML files. But, no: two authoritative Elsevier documents give opposite answers.

Tag by Tag (PDF), the ultimate reference for the Elsevier XML DTDs, is clear. It says that in the case of the second author the given name field contains "A.N." The capture manual used for the secondary databases is equally clear in the opposite direction: there is no given name and "A.N." belongs in the initials field.

<ce:author>
  <ce:given-name>A.N.</ce:given-name>
  <ce:surname>Author</ce:surname>
</ce:author>
<author>
  <ce:initials>A.N.</ce:initials>
  <ce:surname>Author</ce:surname>
</author>

This looks like a documentation error. It is actually more interesting than that.

Younger colleagues may not realise that Elsevier is a stack of acquisitions, each with its own traditions. I come from a department whose traditions were largely inherited from North-Holland, originally a publisher. Even something as innocent as initials required standardization: different publishers we acquired used initials with or without full stops, with or without spaces, and it took years to settle on one convention.

The other bloodline comes from the secondary databases, EMBASE among them, and the capture manual reflects that tradition. The publishing bloodline starts with what appeared in the article; the secondary-database bloodline is more inclined to interpret what those strings mean.

At first glance the capture manual seems more logical. "A.N." is obviously not a given name.

True. But irrelevant.

"Alice N." isn't a given name either – not with that "N." attached – yet according to the same manual that does belong in the given name field if the author had spelled out their first name. And things become much more interesting once we encounter names from naming traditions that do not divide neatly into "given name" and "surname". Can we reliably decide what a given name is? More importantly: why do we need to?

The use case isn't to know how we should address the author at a party. We need to produce a long name ("Alice N. Author"), a space-saving short name ("A.N. Author") and a sort name ("Author, A.N."). And if all the author gives us is "A.N.", then that's the long name.

Regular readers will know my adage: "model the data, not the words". I did not invent this way of thinking. It is the practical lesson handed down by the people who came before me, Fred Veldmeijer in particular. Data modellers don't naturally approach things this way: a field called "given name" creates an almost irresistible urge to decide what really is a given name.

And so the old discussion returns. The conversions for our new generation of products are triggering arguments that we thought we'd settled twenty years ago.

There is a final twist. Although the capture manual reveals the secondary-database bloodline, the data in the citation database itself doesn't. Somewhere along the value chain, those empty given names are populated with the initials, so that given name + surname remains meaningful.

That is what acquisitions leave behind. Systems, manuals and terminology retain traces of different traditions long after one bloodline has become dominant. New colleagues encounter the resulting inconsistencies without knowing the history that produced them.

That history is useful. When two perfectly sensible definitions disagree, the question to ask is not necessarily "Which definition is correct?"

Sometimes the better question is: "What does the data need to do?"

Related posts