Abdolmadjid Masoomi

Metadata Is the Message

Why who you contacted, when, and from where says more than what you said

Published
2026-09-12
Length
4 min read · 693 words
Status
supported not independently verified

Content is expensive to analyse and easy to encrypt. Metadata is cheap to analyse, hard to hide, and sufficient for most conclusions anybody wants to draw about a person. Why the distinction is drawn where it is, and what follows from it.

A worked example

The following is invented, to make a structural point rather than to report anything.

A week of one device's connection log, with no transcripts and no recordings. Monday morning, a four-minute call to a number registered to a health clinic. Tuesday evening, twenty minutes to a close family member. Wednesday, nothing at all. Thursday morning, forty-five minutes to a legal practice. Friday afternoon, several short messages to that same number, then a large file upload.

Not one word of any of it is known, and the week is legible anyway. Something was found; somebody was told; a day was lost to it; a solicitor was engaged and then sent documents. The silence on Wednesday carries as much as any of the calls.

That is the point. The content was fully protected, and the story survived intact.

Why the split exists

The separation between content and metadata is structural, not a policy anyone chose.

A system cannot deliver a message without knowing where it goes. The address has to be legible to every intermediary that handles it, or it does not arrive. Content can be sealed because only the endpoints need to read it; addressing cannot, because the machinery in between has to.

Content is the letter and metadata is the envelope, and the envelope must stay readable to the postal service by the nature of posting things. Encrypting it is not a matter of willingness.

What it contains

Addressing, timing, frequency, duration, size, and the location of the device when it connected.

The powerful part is none of those individually. It is the graph they imply: who communicates with whom, how often, and whom they have in common. One call is an event. Repetition is a relationship. Repetition across a group is a structure, and structures can be searched for.

Pattern of life

Routine is what makes the rest legible. Most people generate a steady, boring baseline, and the baseline is not the interesting part — it is what makes deviation visible.

A change in where a device sleeps. Contact at an hour that has never carried contact before. A new number appearing frequently and then disappearing. None of these requires content to be meaningful, because the meaning is in the departure from an established shape.

This is why pattern analysis scales in a way content analysis does not. Reading everybody's messages is expensive and mostly worthless. Noticing that a pattern changed is cheap and frequently sufficient.

What actually reduces it

Not much, and being honest about that is more useful than a list.

Fewer intermediaries, because each one is a place the envelope is read. Services that do not retain what they do not need, which is a property of the service rather than a setting you can choose. Padding and batching, which genuinely work and are rare precisely because they cost latency that users notice and will not tolerate. Separate identities for separate purposes, which fragments the graph and is effortful enough that almost nobody sustains it.

Individual discipline is weak here in a way it is not for content. This is a design problem. The exposure is produced by the architecture doing its job, and no amount of care by the person using it changes what the architecture must know.

Why "we only collect metadata" is a strange reassurance

The statement is usually true, which is what makes it worth examining rather than dismissing.

It sounds modest because the word sounds clerical — filing, receipts, administrative exhaust. Something kept by a back office rather than something that describes you. The word does a great deal of work in that sentence, and it does it quietly.

But the collection described is the half that yields conclusions. Content is unstructured, various and awkward to process at scale. Metadata is uniform, structured and built for querying. If you were choosing which half to keep in order to understand a population, you would keep this one.

Close

When a service tells you what it collects, read the metadata paragraph first, and read the content paragraph as the part they were comfortable being asked about.