---
title: "SEO"
canonical: "https://www.owox.com/data-models/seo"
updated: "2026-09-26"
---

# SEO Data Model

18 data marts111 fields[Max Roslyakov](https://www.linkedin.com/in/xamsor/)[Vlad Flaks](https://github.com/vladflaks)

A site's content, technical health and backlinks meet where a page ranks for a query on a given day, with crawls, indexing, topical authority and knowledge gain explaining the position and the visits it earns.

## Overview

Organic search as the practice itself describes it. SEO is built on three pillars. The first is content: to be findable by AI and by Google a business has to write about itself on its own site, and what it writes is the material that gets indexed and served. The second is technical: whether the site that hosts it can be crawled, whether robots are told where they may go, whether it is reasonably fast, whether its layout holds up on a phone. The third is off-site — what the rest of the web says about the business. Content, what it is hosted on, and what the outside world says about it. The three meet where a page stands for a query on a given day, and in the arrivals that standing produces.

What makes it more than a funnel is that neither the measurements at the end nor the judgements behind them can be read on their own. A position is not traffic: coming from nowhere into the top 20 is a very noticeable result and brings no traffic yet, getting from there into the first five is another three or four months' work, and 90% of traffic is concentrated in Google's first five lines because people do not scroll down when those five answer them. And the two judgements the model turns on are both made against a topic rather than in the abstract. Authority is a site's standing in a context — standing is topical, and a site can carry it in one subject and none in another — and knowledge gain is what a piece of content brings into a context that other sites have already written on.

The model was drawn from a recorded ontology interview with [Max Roslyakov](https://www.linkedin.com/in/xamsor/), who runs an SEO agency and knows the link market from what his own system sees of it, and it holds what that conversation covered. Rankings are held over time, because the two clocks that move them only show themselves across dates: work done fundamentally — strong content, good links, and no competitor whose own actions start pushing a site off the top places — can stay stable for as much as five years, while a hack stops giving what it was giving once an algorithm update closes it.

**Scope:** the model ends at the visit. There is no lead, order or revenue anywhere in it; it stops at whether an arrival was qualified, because most SEO practitioners do treat the visit as their target event, and everything past it belongs to a different model. It also does not tie a visit to the query that produced it. What an arrival carries is how it arrived, in Google's own terms the channel `organic` with the source `google`, and not the words somebody typed; the search term is what the industry calls `not provided` — a fact about analytics rather than anything this conversation settled — so demand and behaviour meet in the dated aggregate in Search Performance, query by query and page by page, and nowhere else. Competitors are absent as objects: a competitor pushing a site off the top places is part of why a position moves, but neither competitors nor their rankings are held here. Also absent, and worth naming so that nobody looks for them: SERP features — snippets, AI overviews, local packs — internal link structure as a graph, Core Web Vitals broken out one by one, and any ledger of what search work costs. The three things a practice pays for are each here, on the object each belongs to: producing content, developer work to put the site in order, and a placement on somebody else's site. Nothing sums them into a budget. Two absences inside marts that are here are worth the same warning. A keyword carries no search volume, no difficulty and no cost per click, so the model cannot say whether a topic is worth entering at all — the impressions in Search Performance stand in only for the queries a page already appears for. And a backlink records what was bought — whether there is a link at all, whether it was bought as permanent, how many months it is expected to stay up, and the two costs — but not the anchor text, not whether the link is marked nofollow, and not whether the placement is still up today.

## Example Questions

*   Which `queries` have we come from nowhere into the top 20 for, which of those have gone on into the first five lines, and how long did each step take us?
*   For a topic we are trying to enter, which of the sites we pay to appear on carry standing in that topic rather than standing in general, and what have those placements cost us once the articles we had to write for them are counted?
*   Which of our `pages` hold positions and still earn nothing, what does the `content` on them add to its topic, and what are they doing to the averages the rest of the site is read by?

[Explore on canvas →](https://model.owox.com/?okf=https://github.com/OWOX/models/tree/main/bundles/seo)

## Authority

How much standing one external site carries in one topic, on one day. A site's standing is a combination of the organic traffic it already has and the links pointing at it within a context, which is why it is measured per site _and_ per topic and never per site alone: the same site can have standing on web hosting and none at all on washing machines. Two rows for the same site in two topics are not a duplicate — they are the point.

Neither ingredient is invented here. The traffic is held on the [host](#mart-third-party-websites), because it is true of the site whatever topic is being asked about; the links are what makes the reading topical. One exception is real and is kept on the host as its own flag: the broadly authoritative sites that count in any context.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `authority_id` | STRING | Authority ID | PK. The identifier of this reading — one site, in one topic, on one day. |
| `third_party_website_id` | STRING | Third Party Website ID | The external site whose standing is being measured. FK to [Third Party Websites](#mart-third-party-websites) |
| `context_id` | STRING | Context ID | The topic the standing is measured in. The same site can hold standing in one topic and none in another. FK to [Context](#mart-context) |
| `authority_score` | FLOAT | Authority Score | How much standing the site carries in this topic — a combination of the organic traffic it already has and the links pointing at it within this context. |
| `measured_on` | DATE | Measured On | The day this reading was taken. A score is comparable with another only alongside its date. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Third Party Websites](#mart-third-party-websites) | `third_party_website_id = third_party_website_id` | N:1 | The site whose standing is being scored. |
| [Context](#mart-context) | `context_id = context_id` | N:1 | The topic it is scored in; a site can be authoritative in one and unknown in another. |

## Backlink

One placement of a brand on somebody else's site: a link back to one of its pages, or the brand named in their text without one. Both are held here because both are the same piece of work — a business appears in someone else's article, in some topic, on a host whose standing is what decides the worth. `has_link` tells the two apart, and false is not a missing value: it is a brand mention, and a mention points at no page, so a count that reads an empty page as a defect is counting mentions as broken links.

What decides what a placement is worth is the topic it appears in. Standing is topical: a host can carry it in one subject and none in another, and a placement whose article sits outside the subject its host carries weight in is worth quite moderately in Google's eyes. The [broadly authoritative sites](#mart-third-party-websites) are the exception: a piece on Forbes mentioning a brand counts heavily in its favour in any context.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `backlink_id` | STRING | Backlink ID | PK. The identifier of this placement, whether it carries a link or is a mention without one. |
| `page_id` | STRING | Page ID | The page on this site the link points at. Empty when `has_link` is false — a mention points at no page. FK to [Page](#mart-page) |
| `third_party_website_id` | STRING | Third Party Website ID | The external site the placement sits on. FK to [Third Party Websites](#mart-third-party-websites) |
| `context_id` | STRING | Context ID | The topic the placement appears in, which is what decides what it is worth. FK to [Context](#mart-context) |
| `has_link` | BOOLEAN | Has Link | Whether the placement carries a link back to this site. False is a brand mention: the brand named in the text, not linked — the signal AI search responds to. |
| `placed_on` | DATE | Placed On | When the placement went live on the other site. |
| `is_permanent` | BOOLEAN | Is Permanent | Whether it was bought as a permanent placement — the trade's own word for one that lives on the site permanently, rather than a subscription. |
| `min_placement_months` | INTEGER | Minimum Placement Months | How long the placement is expected to stay up at the least. Twelve months is the threshold usually expected. |
| `placement_cost` | NUMERIC | Placement Cost | What was paid to the owner of the site for carrying the placement. |
| `content_cost` | NUMERIC | Content Cost | What it cost to produce the article that carries the placement — the second half of the price, paid to whoever wrote it rather than to the site. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Page](#mart-page) | `page_id = page_id` | N:1 | The page the link points at; empty when the placement is a mention with no link. |
| [Third Party Websites](#mart-third-party-websites) | `third_party_website_id = third_party_website_id` | N:1 | The site the placement sits on. |
| [Context](#mart-context) | `context_id = context_id` | N:1 | The topic the placement appears in, which is what decides whether it counts. |

## Channel

How a visitor arrived, as an analytics tool classifies it: organic search, a referral from another site, direct entry, or an AI assistant. In Google's own terms a visit from search is the channel `organic` with the source `google`: the channel says what kind of arrival it was, the source says who sent it.

Search work moves more than one of these at once, which is why the classification is kept as its own object rather than as a label on the visit. A placement on a good site is read by search engines as a citation, and people follow it: those arrivals are `referral`, while the ranking the same placement supports produces `organic` ones. The same piece of work arrives under two labels, and a practice reading only their sum cannot separate the two effects.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `channel_id` | STRING | Channel ID | PK. The identifier of this way of arriving. |
| `channel_group` | STRING | Channel Group | The kind of arrival: `organic`, `referral`, `direct` or `ai`. |
| `source` | STRING | Source | Who sent the visitor — `google` for organic search from Google, the domain for a referral. |

## Content

The text, images and video a site publishes: the first of the three pillars of search — what a business says about itself, and the material that gets indexed and served in AI answers and in Google. It is also one of the three things an SEO practice pays for, alongside the developer work that keeps the site in order and the placements bought on other people's sites, so this is where the first of those three cost lines sits.

Reading content as a volume is the reading this model is built to prevent. What gives a piece of content a chance of ranking is not that it exists, and not that a person rather than a machine produced it, but whether it adds knowledge to its topic. Content that adds nothing still costs what it cost to produce; and a large body of generated content that does not rank demonstrably worsens a site's traffic figures, often looking like a collapse.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `content_id` | STRING | Content ID | PK. The identifier of this piece of content. |
| `page_id` | STRING | Page ID | The page this text, image or video sits on, and the only route from content to the site that publishes it. FK to [Page](#mart-page) |
| `knowledge_gain_id` | STRING | Knowledge Gain ID | The assessment of what new knowledge this content brings into its topic. FK to [Knowledge Gain](#mart-knowledge-gain) |
| `content_type` | STRING | Content Type | What kind of content it is: `text`, `image` or `video`. |
| `published_at` | TIMESTAMP | Published At | When this content went live on the page. |
| `is_ai_generated` | BOOLEAN | Is AI Generated | Whether the content was produced by AI. On its own this decides nothing: what matters is whether the content adds knowledge to its topic. |
| `production_cost` | NUMERIC | Production Cost | What it cost to create this content — one of the three things an SEO practice pays for, alongside developer work on the site and placements on other sites. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Page](#mart-page) | `page_id = page_id` | N:1 | The page this text, image or video sits on. |
| [Knowledge Gain](#mart-knowledge-gain) | `knowledge_gain_id = knowledge_gain_id` | N:1 | The assessment of what new knowledge this content adds. |

## Context

The topic a search belongs to — the subject someone is asking about, rather than the words they used to ask. Many different queries and many different keywords can be the same context: "web analytics for websites", "visit analytics for e-commerce" and "how do I find out where my customers came from" are all one topic, and a practice works on the topic rather than on the phrasings one at a time.

It is a mart of its own, with almost nothing on it, because it is what the rest of the model is judged against. Authority is contextual — standing is topical, and a site can carry it in one subject and none in another — and whether a page adds new knowledge is measured against a topic. Leaving the context out is what an entry-level practitioner does: calling a site authoritative because it is large and gets traffic, and buying a placement on it whose value in Google's eyes turns out to be quite moderate. The more often a site is cited within a context, the better its chance of coming up, and coming up high, for it.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `context_id` | STRING | Context ID | PK. The identifier of this topic. |
| `topic` | STRING | Topic | The subject itself, in the practice's own words — the thing a set of queries, a set of keywords and a body of content are all about. |

## Crawl

One fetch of one page by one robot. A row is a single arrival, not a robot: a robot can fetch the same page more than once, and each fetch is its own row. Keeping the grain at the fetch is what lets crawling be read as behaviour over time rather than as a list of who is out there.

A crawler landing on a site learns its content and its links. The internal ones lead it to internal pages, which it also visits — each of those visits another row here. An external one may send it to a page it did not know, and that page belongs to somebody else and has no row in this model, so the fetch names a [host](#mart-third-party-websites) and no page.

That is why this mart reaches the [crawl rules](#mart-robots-txt) by two routes rather than one. A fetch of the site's own page reaches them through the page and its [site](#mart-website); a fetch of somebody else's page has only the route through the [indexing request](#mart-initiate-index) — one is the rules of the site whose page was fetched, the other the rules the run itself read.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `crawl_id` | STRING | Crawl ID | PK. The identifier of this single fetch. |
| `page_id` | STRING | Page ID | The page that was fetched, when the fetch was of this site. Empty when the crawler followed a link off it onto somebody else's page. FK to [Page](#mart-page) |
| `third_party_website_id` | STRING | Third Party Website ID | The external host the robot went to, when the page fetched belonged to somebody else. FK to [Third Party Websites](#mart-third-party-websites) |
| `initiate_index_id` | STRING | Indexing Request ID | The run that caused this fetch, which is how a request is tied to the pages it actually reached. FK to [Initiate Index](#mart-initiate-index) |
| `crawler_user_agent` | STRING | Crawler User Agent | Who fetched the page, as it introduced itself. Anyone requesting a page on the internet is obliged to state a user agent, and usually that is more than enough to tell a robot from a person. |
| `crawled_at` | TIMESTAMP | Crawled At | When this fetch happened. |
| `was_allowed` | BOOLEAN | Was Allowed | The crawl rules verdict on this fetch: whether the robot was let into the path it asked for. |
| `links_found` | INTEGER | Links Found | How many links the crawler came away with. Internal ones lead it to pages it then visits too; external ones either credit a page it already knew or send it to crawl one it did not. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Page](#mart-page) | `page_id = page_id` | N:1 | The page fetched, when the fetch was of this site; empty for a third-party page. |
| [Third Party Websites](#mart-third-party-websites) | `third_party_website_id = third_party_website_id` | N:1 | The external site fetched, when the crawler followed a link off this one. |
| [Initiate Index](#mart-initiate-index) | `initiate_index_id = initiate_index_id` | N:1 | The indexing run this fetch belongs to. |

## Google Search Console

The property registered with Google: Google's own facility where a site is entered directly, as a claim of ownership and a request that it be crawled — and Google goes and crawls it. It is the one place in this model where a site is declared to a search engine instead of being discovered by one, which is why it is held as an object of its own rather than as a flag on the site.

It carries no numbers. Impressions, clicks and positions are read _through_ this property and are kept in [Search Performance](#mart-search-performance); what lives here is the registration itself — which site, whether the claim was confirmed, and when. Search Console is also where searching done by AI is fairly visible, as very long queries a normal person would never write.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `property_id` | STRING | Property ID | PK. The identifier of this registered property. |
| `website_id` | STRING | Website ID | The site this property was registered for. FK to [Website](#mart-website) |
| `property_url` | STRING | Property URL | The address the property covers, as it was entered when the site was registered. |
| `is_verified` | BOOLEAN | Is Verified | Whether ownership of the site has been confirmed, which is what makes the property a claim rather than an entry. |
| `verified_at` | TIMESTAMP | Verified At | When ownership was confirmed, so a change in crawling or in search numbers can be lined up against the moment the site was declared. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Website](#mart-website) | `website_id = website_id` | N:1 | The site this property was verified for. |

## Initiate Index

The moment a site or one of its pages enters the queue to be indexed. The name is the one coined on the call this model was drawn from rather than a phrase used elsewhere in the trade.

A run arrives here by one of three routes. A site can be entered into Google's [Search Console](#mart-google-search-console) directly, as a claim of ownership and a request that it be crawled. A [sitemap](#mart-sitemap-xml) can offer the structured list of pages to be indexed. Or a crawler already working somewhere else can find a link to a page it did not know, and go and crawl that page. The sitemap route is optional: a sitemap helps, but if there is none, Google will crawl the site anyway, as long as the robots.txt is in order.

The date here is when the asking happened, not when a robot turned up, and the gap to the [fetches](#mart-crawl) that followed is what shows how long a site waited.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `initiate_index_id` | STRING | Indexing Request ID | PK. The identifier of this run. |
| `robots_txt_id` | STRING | Robots File ID | The crawl rules the arriving crawler reads on the way in, whichever route asked for the run. FK to [robots.txt](#mart-robots-txt) |
| `requested_at` | TIMESTAMP | Requested At | When the site or page was put into the queue, which is not when a robot arrived. |
| `trigger` | STRING | Trigger | Which route started the run: `search_console` for a request made directly in the registered property, `sitemap` for the page list being offered, `external_link` for a crawler that met a link to a page it did not know. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [robots.txt](#mart-robots-txt) | `robots_txt_id = robots_txt_id` | N:1 | The rules the run read before fetching anything. |

## Keyword

A term the practice targets — the middle of the three levels this model keeps for demand. It is one of the main terms of the trade. A keyword is, to a considerable extent, a synonym of a [topic](#mart-context): many keywords can be one and the same context, and the practice works on the topic rather than on each phrasing separately.

The word carries two meanings in everyday use, and this model splits them. Asked directly whether the thing a user sees is a keyword or a search query, the answer on the record was: it is a search query. So the string somebody typed is a [Search Query](#mart-search-query), and `keyword` is used here by this model's own convention for the term the practice targets — the entry on the practice's list, not the entry in somebody's search box. Holding the two apart is what makes it possible to ask which of the chosen terms nobody searches, and which searches a site picks up that it never aimed at.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `keyword_id` | STRING | Keyword ID | PK. The identifier of this term. |
| `context_id` | STRING | Context ID | The topic this term belongs to; a keyword is to a considerable extent a synonym of a context, and many keywords can be one and the same one. FK to [Context](#mart-context) |
| `keyword_text` | STRING | Keyword | The term itself, as the practice writes it on its own list. |
| `is_targeted` | BOOLEAN | Is Targeted | Whether the practice is optimising for this term, as opposed to keeping it on the list to watch. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Context](#mart-context) | `context_id = context_id` | N:1 | The topic this term belongs to; many keywords share one. |

## Knowledge Gain

Whether a piece of content brings new knowledge into a topic — a super-important component of ranking. Google holds patented technology for the judgement, published information rather than a secret, by which it assesses how valuable a page is in bringing new knowledge into a context. It is kept as an object of its own rather than as a column on the content because there is a context, there is content other sites have already written on that context, and what is being measured is the difference a page makes to it.

It is the brick people fail to understand when they hand content creation over to AI and are then surprised at how their site is doing. The term is not much circulated, and it stays uncirculated because it gets in the way of selling AI content generators. It is not a verdict on the author: content that merely retells what is already available carries no gain, and content built on data nobody else holds carries it whether or not a machine wrote the words.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `knowledge_gain_id` | STRING | Knowledge Gain ID | PK. The identifier of this assessment. |
| `context_id` | STRING | Context ID | The topic the content is being judged against — there is no such thing as a gain in the abstract. FK to [Context](#mart-context) |
| `score` | FLOAT | Knowledge Gain Score | How much new knowledge this content brings into its topic, as a number. It is read against the topic, not against the content on its own. |
| `adds_new_knowledge` | BOOLEAN | Adds New Knowledge | Whether the content brings knowledge that was not already in general circulation, or only retells what other sites have written on the topic. |
| `is_expert_authored` | BOOLEAN | Is Expert Authored | Whether the author is already visible in this topic as the author of content, ideas, opinions and analysis — which gives a page a chance of ranking despite weak support from links. |
| `is_based_on_unique_data` | BOOLEAN | Is Based on Unique Data | Whether the content rests on data that was not otherwise available. Where it does, machine authorship makes no difference to how it ranks. |
| `contradicts_consensus` | BOOLEAN | Contradicts Consensus | Whether the content argues against an established consensus in its topic. Content that does is deranked even when it is right. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Context](#mart-context) | `context_id = context_id` | N:1 | The topic the gain is measured against; the same content scores differently in another. |

## Page

One page of the site being optimised: the unit that gets crawled, indexed, ranked and arrived at. A page is what content sits on, and the level at which what a site asks a search engine for — index this — can be compared with what happened to it. `is_in_sitemap` is a column here and not a link to the [page list](#mart-sitemap-xml): the list is optional, and being named in it is simply true or false about a page.

It carries its own behaviour numbers because Google reads user signals per page _and_ for the site as a whole, as an average. A hack turns on that arithmetic: removing content that does not work lifts the averages of what is left, which Google reads as an improvement in quality. Whether a dead page should be removed, redirected or rewritten is not settled, and depends heavily on how spoilt Google is for content on that topic: where it has plenty, such pages should not be on the site at all; where it has little, it will show even weak ones for want of anything better. Pages kept for regulatory or documentation reasons can be held quite safely.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `page_id` | STRING | Page ID | PK. The identifier of this page. |
| `website_id` | STRING | Website ID | The site this page belongs to, which is what separates the pages of the site being optimised from anyone else's. FK to [Website](#mart-website) |
| `url` | STRING | Page URL | Where the page is published. |
| `page_type` | STRING | Page Type | What kind of page it is: `home`, `service`, `contact`, `case_study` or `blog`. |
| `is_indexable` | BOOLEAN | Is Indexable | Whether the page is telling search engines it wants to be indexed. |
| `is_indexed` | BOOLEAN | Is Indexed | Whether the page is actually in the index. It can differ from what the page asks for. |
| `is_in_sitemap` | BOOLEAN | Is In Sitemap | Whether the site's page list names this page. The list itself is optional — a site with none is still crawled. |
| `last_crawled_at` | TIMESTAMP | Last Crawled At | When a search robot last fetched this page. |
| `avg_time_on_page` | FLOAT | Average Time on Page | How long visitors stay on this page — a user signal Google reads per page as well as averaged over the site. |
| `avg_depth_of_scroll` | FLOAT | Average Depth of Scroll | How far down this page visitors get before they stop. |
| `bounce_rate` | FLOAT | Bounce Rate | The share of visitors who arrive on this page and leave straight away. |
| `has_organic_traffic` | BOOLEAN | Has Organic Traffic | Whether the page earns any traffic from search at all. Pages that earn none are the ones a removal experiment is aimed at, and the ones that move the site's averages down. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Website](#mart-website) | `website_id = website_id` | N:1 | The site this page belongs to, which is what separates the pages of the site being optimised from anyone else's. |

## robots.txt

The file a site publishes to tell arriving search robots where they may go and where they may not, and which of them the owner wants indexed at all. It is its own object rather than a field on the site because it decides what everything downstream is even allowed to see.

A site that publishes no sitemap is still crawled, as long as its robots.txt is in order. A path closed here is a path the crawler does not fetch, whatever the page behind it says about wanting to be indexed — which is why the rules and the pages are worth reading together rather than separately. The count of closed paths says nothing on its own about whether the closures are right: a site that deliberately keeps part of itself out of the index has a high count and a sound setup, and a site that has closed a content directory by mistake looks identical from here.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `robots_txt_id` | STRING | Robots File ID | PK. The identifier of this crawl-rules file. |
| `url` | STRING | File URL | Where the file is published, at the root of the site it governs. |
| `is_valid` | BOOLEAN | Is Valid | Whether the file parses as valid crawl rules. When it does not, an arriving crawler is left without the instructions it came for. |
| `disallowed_path_count` | INTEGER | Disallowed Paths | How many paths the file closes to crawlers. Read against what the pages behind them are meant to do, not on its own. |
| `last_modified_at` | TIMESTAMP | Last Modified At | When the rules last changed, so a shift in crawl behaviour can be lined up against a change in what the crawler was told. |

## Search Performance

Where one page stood for one query on one day, and what that standing earned. A page can hold quite different positions for two queries on the same day, so the key is all three.

The distinction this model turns on lives here: a position is not traffic. Coming from nowhere into the top 20 is a very noticeable result, and there is no traffic yet. Getting from the top 20 into the top five is another three or four months of work: 90% of traffic is concentrated in Google's first five lines, because people do not scroll down when the first five answer them. Searches made by AI land here beside the ones people typed, and a position averaged over both puts the two kinds into one figure.

Because every row is dated, the two clocks of SEO can be told apart here. Work done fundamentally — strong content, good links, and no competitor whose own work starts pushing a site off the top places — can stay stable for as much as five years. A hack can be closed off by an algorithm update, after which it stops giving what it was giving.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `date` | DATE | Date | PK. The day these figures were measured for, which is what makes a position readable as movement rather than as a single snapshot. |
| `search_query_id` | STRING | Search Query ID | PK. The query this row's impressions and position were measured for. FK to [Search Query](#mart-search-query) |
| `page_id` | STRING | Page ID | PK. The page that appeared in the results for that query. FK to [Page](#mart-page) |
| `impressions` | INTEGER | Impressions | How many times the page was shown in the results for this query on this day. |
| `clicks` | INTEGER | Clicks | How many times someone went from the results to the page. |
| `ctr` | FLOAT | Click-Through Rate | The share of impressions that became clicks. |
| `avg_position` | FLOAT | Average Position | Where the page stood in the results that day, averaged over its impressions. It is not traffic: a page in the top 20 has a real result and no visits yet, while 90% of traffic sits in the first five lines. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Search Query](#mart-search-query) | `search_query_id = search_query_id` | N:1 | The query this row's impressions and position were measured for. |
| [Page](#mart-page) | `page_id = page_id` | N:1 | The page that appeared in the results for it. |

## Search Query

What somebody actually put into a search box. It is the finest of the three levels this model keeps for demand: a query is the string that was typed, a [keyword](#mart-keyword) is — by this model's own convention — the term the practice targets, and a [context](#mart-context) is the topic both of them sit in.

The level is where ranking is settled: a position is a position _for a query_. In the practitioner's own picture of it, for a given query a site was not taking part in the lottery at all, and now it is in it and holding a place.

The boundary between a query and a keyword has blurred a little. A search made by AI is, as a rule, not a keyword but a pile of key words — a very long string a normal person would never write — and those searches land in this mart next to the short ones a person typed.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `search_query_id` | STRING | Search Query ID | PK. The identifier of this query string. |
| `keyword_id` | STRING | Keyword ID | The term the practice targets that this typed query was matched to. FK to [Keyword](#mart-keyword) |
| `query_text` | STRING | Search Query | The string that was actually searched for. |
| `is_ai_generated_query` | BOOLEAN | Is AI-Generated Query | Whether the query looks like a search made by AI rather than typed by a person: as a rule not a keyword but a pile of key words, a very long string a normal person would never write. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Keyword](#mart-keyword) | `keyword_id = keyword_id` | N:1 | The term the practice targets that this typed query was matched to. |

## sitemap.xml

The structured list of its own pages a site hands to search engines for indexing. It is the site saying, in one place, what it consists of — which is help a crawler can use, but not help it depends on: removing the sitemap does not worsen a site's results in search, as long as everything else is in order.

A site may publish none at all, and the absence is a legitimate state rather than a gap in the data, which is why the list is held separately rather than folded into the site record. Where a sitemap does exist, the number it carries is a claim about what the site contains, and a claim that has drifted away from the pages actually published is worth knowing about. Publishing [content](#mart-content) and telling search engines about it are two separate acts, and the gap between them is visible here.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `sitemap_id` | STRING | Sitemap ID | PK. The identifier of this page list. |
| `url` | STRING | Sitemap URL | Where the list is published for search engines to fetch. |
| `listed_page_count` | INTEGER | Listed Pages | How many pages the list offers for indexing — the site's own claim about what it consists of, to be read against the pages it actually publishes. |
| `last_submitted_at` | TIMESTAMP | Last Submitted At | When the list was last handed to search engines, which is a separate act from publishing the pages in it. |

## Third Party Websites

The sites a practice does not own: the hosts that link to it, mention it, or could. Off-site work is the third pillar of search — what the rest of the web says about a business.

It deliberately carries no single standing for the site: standing combines the organic traffic a host already has with the links pointing at it within a topic, so it is measured per site _and_ per topic and kept on [its own object](#mart-authority). What lives here is what is true of the host whatever the topic — its traffic, and whether it is one of the broadly authoritative sites, Forbes and Wikipedia, where being mentioned in any context counts heavily in a site's favour.

Two things have made this work much more complicated than it used to be. Buying links by the hundred used to be the whole game, and that is no longer how it works; reputation and topical relevance are what count now. And AI does not crawl sites, at least not yet: what it responds to is a brand being mentioned in a context, and, judging by indirect signals, the reputation of the source matters a great deal.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `third_party_website_id` | STRING | Third Party Website ID | PK. The identifier of this external host. |
| `domain` | STRING | Domain | The host itself, as it appears on the placements that sit on it. |
| `organic_traffic` | INTEGER | Organic Traffic | The search traffic the host already receives — one of the two ingredients behind its standing, the other being the links it holds within a topic. |
| `is_general_authority` | BOOLEAN | Is General Authority | Whether this is one of the broadly authoritative hosts that count in any context, rather than in one topic. |

## Visit

Somebody — or something — arriving on one of a site's pages. This is the event the whole model points at: most practitioners do treat the visit as their target event, and it is where this model stops. What happens past the arrival belongs to a different model.

There is no key from a visit to the query that produced it, and that is deliberate rather than missing. What an arrival brings with it is how it arrived — the channel and the source — and not the words somebody typed to get there. Demand and behaviour therefore meet in the daily aggregate in [Search Performance](#mart-search-performance), which is keyed by query and page, and nowhere else.

A visit is also the raw material of the user signals that decide whether a position survives: of the sites somebody opens for one query, the one they spent three minutes on and the one they left at once are read differently, and the one they left at once is thrown out of the top five so the next site can have its chance.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `visit_id` | STRING | Visit ID | PK. The identifier of this arrival. |
| `page_id` | STRING | Page ID | The page that was landed on. FK to [Page](#mart-page) |
| `channel_id` | STRING | Channel ID | How the visitor arrived, as an analytics tool classifies it. FK to [Channel](#mart-channel) |
| `occurred_at` | TIMESTAMP | Occurred At | When the arrival happened. |
| `user_agent` | STRING | User Agent | How the visitor introduced itself. Anyone requesting a page on the internet is obliged to state a user agent. |
| `is_bot` | BOOLEAN | Is Bot | Whether this was a robot rather than a person, read off the user agent — usually more than enough to tell the two apart. |
| `pages_viewed` | INTEGER | Pages Viewed | How many pages the visitor opened on the site during this arrival. |
| `duration_seconds` | INTEGER | Duration in Seconds | How long the arrival lasted. |
| `bounced` | BOOLEAN | Bounced | Whether the visitor left straight away instead of going any further into the site. |
| `is_qualified` | BOOLEAN | Is Qualified | Whether this arrival met the practice's own definition of a visit worth having. It is the last thing the model says about a visit. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [Page](#mart-page) | `page_id = page_id` | N:1 | The page the visit landed on. |
| [Channel](#mart-channel) | `channel_id = channel_id` | N:1 | How the visitor arrived. |

## Website

The site being optimised — the object the whole technical side of search work is done to, and the boundary between what is within a practice's control and what is not. Content is placed on it, and its crawl rules and page list belong to it; the part of SEO that concerns a site's own condition ends where this record ends.

Google sees how people behaved after arriving, and it looks at those user signals both per page and for the site overall, as an average — so speed, layout and the behaviour numbers are properties of the site rather than of any one page. A site need not be superfast, but it does have to be reasonably fast, and 80% of its traffic may be coming from phones while its layout breaks on them. The two connect: a visitor who opens the site on a phone and leaves at once, because the content they came for could not be read, earns the site a minus for user signals, and the site slides down in search.

### Fields

| Column | Type | Alias | Description |
| --- | --- | --- | --- |
| `website_id` | STRING | Website ID | PK. The identifier of this site. |
| `domain` | STRING | Domain | The domain the site is published under — the name everything else in the model is claimed and measured against. |
| `robots_txt_id` | STRING | Robots File ID | The crawl-rules file published at the root of this site, which tells an arriving robot where it may go. FK to [robots.txt](#mart-robots-txt) |
| `sitemap_id` | STRING | Sitemap ID | The list of its own pages this site offers for indexing. Empty when the site publishes none, which does not stop the site being crawled, as long as its robots.txt is in order. FK to [sitemap.xml](#mart-sitemap-xml) |
| `page_speed_score` | FLOAT | Page Speed Score | How quickly the site loads. The bar is not superfast — it is reasonably fast. |
| `mobile_optimization_score` | FLOAT | Mobile Optimization Score | How well the layout holds up across device types, read against how much of the site's traffic arrives from phones. |
| `avg_time_on_page` | FLOAT | Average Time on Page | Time on page averaged across the site — one of the user signals Google reads for the site as a whole and not only per page. |
| `avg_depth_of_scroll` | FLOAT | Average Depth of Scroll | How far down a page visitors get, averaged across the site. |
| `bounce_rate` | FLOAT | Bounce Rate | The share of arrivals that leave straight away, whatever sent them away — being unable to read the content on a phone is one such reason. |
| `development_cost` | NUMERIC | Development Cost | What has been spent on developer work to put the site in technical order and keep it there. |

### Relationships

| Related data mart | On | Cardinality | Meaning |
| --- | --- | --- | --- |
| [robots.txt](#mart-robots-txt) | `robots_txt_id = robots_txt_id` | N:1 | The file that tells crawlers where on this site they may go. |
| [sitemap.xml](#mart-sitemap-xml) | `sitemap_id = sitemap_id` | N:1 | The optional index of the site's pages; empty when the site publishes none. |

## Apply to your project

1.  1

    ### Install the Import Model plugin

    One plugin, installed once, in your own OWOX workspace.

    [Get the plugin →](https://github.com/OWOX/import-model)

2.  2

    ### Import this model

    Point it at this bundle and it creates every data mart above, joins and all.

    [Open the model →](https://model.owox.com/?okf=https://github.com/OWOX/models/tree/main/bundles/seo)

3.  3

    ### Plug in your data and destinations

    Connect your own sources and send the results where your team already works.

    [Browse connectors →](/connectors)

**4\. Optional — customize as you wish.** Rename a column, drop a mart, add your own: once it is imported it is yours, and nothing here syncs back.

## References

Pages this page links to, on this site and on docs.owox.com. Where the page has a Markdown twin, its address follows the link.

- [Browse connectors →](https://www.owox.com/connectors) — /connectors.md
