Appearance
Content
A content item is one piece of creative work we hold: a post, an advert, a page, a film. It is the atomic unit of the library — the thing you look at, save, compare, and put on a board. Almost everything else in the model either describes a content item, gathers a set of them, or points at one.
The test for whether something is a content item is simple: could a person open it and look at it as a single piece of work? A film is one. A landing page is one. A campaign is not — it is a container for several.
How content connects to everything else
Read it as sentences: a company makes content and appears in it; a campaign gathers content together; a video is cut into moments; and every piece is classified against our shared vocabulary.
The kinds of content
Every piece is one of five kinds. Three are in use today; audio and interactive are recognised but empty.
| Kind | What it is |
|---|---|
| Image | A still: a social post, an advert, a photograph, a screenshot of a page |
| Text | A written piece: an article, a press release, a blog post, a paper |
| Video | A film, a spot, a clip, a screen recording |
| Audio | Recognised, nothing held yet |
| Interactive | Recognised, nothing held yet |
Where each piece came from
Every piece records how we collected it — a landing page, an Instagram post, a Meta or LinkedIn advert, a YouTube video, a newsroom feed. This turns out to be the most reliable thing we know about it, more reliable than anything a model has said about it.
From that, a rule works out the channel: the kind of place the work appeared — a website, social, a paid advert, video, a blog, a newsroom.
This is worth being precise about, because it is easy to get backwards. The channel is not a judgement a model makes — it is derived, automatically, from how the piece was collected. A page pulled from a company's newsroom is newsroom content because of where it came from, not because something decided it read like news. That was a deliberate change: the source knows better than the guess.
A channel is a lens, not a container. This matters for how the whole model hangs together. Instagram, TikTok and YouTube are not boxes that own content; they are answers to the question where did this appear? — the same kind of answer as what industry is this? or who is it aimed at? You classify a piece against a channel; you do not put it inside one. The things content genuinely sits inside — campaigns, events, projects — are a separate family, and they are things that actually happened in the world.
What we say about a piece
Beyond where it came from, a piece accumulates description over time. Some of it is a single value on the piece itself; most of it is a list of labels attached alongside.
| What we record | What it answers |
|---|---|
| The company or person it is about | Whose work is this? |
| The themes it covers | What is it talking about? |
| The crafts involved in making it | What skills went into it? |
| Its written text or transcript | What does it actually say? |
| What kind of thing it is | Is it an advert, a case study, a brand film? |
| The creative concept behind it | What is the idea? |
| The campaign it belongs to | What story is it part of? |
These fill in very unevenly, and the shape of that unevenness is the real story of this pillar. What we can say confidently is where a piece came from and who it is about. What we can say about what it is and how it was made thins out fast.
A real example, straight from our data
OpenAI is the brand almost all of this library is about. We hold 12,790 pieces of its work — 91.7% of everything we have. By kind: image 7,467 · text 3,042 · video 2,281. By where it appeared: website 4,465 · social 3,505 · paid ad 2,910 · blog 813 · video 600 · newsroom 485.
That single number is the most important thing on this page. The model described here has really only been tested against one brand's work, in one industry. Everything looks tidy at this size; the questions that matter are the ones a second and third brand will ask of it.
Where to look next
- Every value you can actually choose — the full vocabulary, one page per axis.
- How much of each of these is actually filled in — the state of the data.
- Where the live system disagrees with this page — the conformance report.
- What is still undecided about content — the open questions.
- The exact tables and columns — the data reference.