SEO Data Model

Overview

Organic search as the practice itself describes it. SEO is built on three pillars. The first is content: to be findable by AI and by Google a business has to write about itself on its own site, and what it writes is the material that gets indexed and served. The second is technical: whether the site that hosts it can be crawled, whether robots are told where they may go, whether it is reasonably fast, whether its layout holds up on a phone. The third is off-site — what the rest of the web says about the business. Content, what it is hosted on, and what the outside world says about it. The three meet where a page stands for a query on a given day, and in the arrivals that standing produces.

What makes it more than a funnel is that neither the measurements at the end nor the judgements behind them can be read on their own. A position is not traffic: coming from nowhere into the top 20 is a very noticeable result and brings no traffic yet, getting from there into the first five is another three or four months' work, and 90% of traffic is concentrated in Google's first five lines because people do not scroll down when those five answer them. And the two judgements the model turns on are both made against a topic rather than in the abstract. Authority is a site's standing in a context — standing is topical, and a site can carry it in one subject and none in another — and knowledge gain is what a piece of content brings into a context that other sites have already written on.

The model was drawn from a recorded ontology interview with Max Roslyakov, who runs an SEO agency and knows the link market from what his own system sees of it, and it holds what that conversation covered. Rankings are held over time, because the two clocks that move them only show themselves across dates: work done fundamentally — strong content, good links, and no competitor whose own actions start pushing a site off the top places — can stay stable for as much as five years, while a hack stops giving what it was giving once an algorithm update closes it.

Scope: the model ends at the visit. There is no lead, order or revenue anywhere in it; it stops at whether an arrival was qualified, because most SEO practitioners do treat the visit as their target event, and everything past it belongs to a different model. It also does not tie a visit to the query that produced it. What an arrival carries is how it arrived, in Google's own terms the channel organic with the source google, and not the words somebody typed; the search term is what the industry calls not provided — a fact about analytics rather than anything this conversation settled — so demand and behaviour meet in the dated aggregate in Search Performance, query by query and page by page, and nowhere else. Competitors are absent as objects: a competitor pushing a site off the top places is part of why a position moves, but neither competitors nor their rankings are held here. Also absent, and worth naming so that nobody looks for them: SERP features — snippets, AI overviews, local packs — internal link structure as a graph, Core Web Vitals broken out one by one, and any ledger of what search work costs. The three things a practice pays for are each here, on the object each belongs to: producing content, developer work to put the site in order, and a placement on somebody else's site. Nothing sums them into a budget. Two absences inside marts that are here are worth the same warning. A keyword carries no search volume, no difficulty and no cost per click, so the model cannot say whether a topic is worth entering at all — the impressions in Search Performance stand in only for the queries a page already appears for. And a backlink records what was bought — whether there is a link at all, whether it was bought as permanent, how many months it is expected to stay up, and the two costs — but not the anchor text, not whether the link is marked nofollow, and not whether the placement is still up today.

Example Questions

  • Which queries have we come from nowhere into the top 20 for, which of those have gone on into the first five lines, and how long did each step take us?
  • For a topic we are trying to enter, which of the sites we pay to appear on carry standing in that topic rather than standing in general, and what have those placements cost us once the articles we had to write for them are counted?
  • Which of our pages hold positions and still earn nothing, what does the content on them add to its topic, and what are they doing to the averages the rest of the site is read by?

Explore on canvas →

Authority Data Mart

Authority

How much standing one external site carries in one topic, on one day. A site's standing is a combination of the organic traffic it already has and the links pointing at it within a context, which is why it is measured per site and per topic and never per site alone: the same site can have standing on web hosting and none at all on washing machines. Two rows for the same site in two topics are not a duplicate — they are the point.

Neither ingredient is invented here. The traffic is held on the host, because it is true of the site whatever topic is being asked about; the links are what makes the reading topical. One exception is real and is kept on the host as its own flag: the broadly authoritative sites that count in any context.

Fields

ColumnTypeAliasDescription
authority_idSTRINGAuthority IDPK. The identifier of this reading — one site, in one topic, on one day.
third_party_website_idSTRINGThird Party Website IDThe external site whose standing is being measured. FK to Third Party Websites
context_idSTRINGContext IDThe topic the standing is measured in. The same site can hold standing in one topic and none in another. FK to Context
authority_scoreFLOATAuthority ScoreHow much standing the site carries in this topic — a combination of the organic traffic it already has and the links pointing at it within this context.
measured_onDATEMeasured OnThe day this reading was taken. A score is comparable with another only alongside its date.

Relationships

Related data martOnCardinalityMeaning
Third Party Websitesthird_party_website_id = third_party_website_idN:1The site whose standing is being scored.
Contextcontext_id = context_idN:1The topic it is scored in; a site can be authoritative in one and unknown in another.

Channel Data Mart

Channel

How a visitor arrived, as an analytics tool classifies it: organic search, a referral from another site, direct entry, or an AI assistant. In Google's own terms a visit from search is the channel organic with the source google: the channel says what kind of arrival it was, the source says who sent it.

Search work moves more than one of these at once, which is why the classification is kept as its own object rather than as a label on the visit. A placement on a good site is read by search engines as a citation, and people follow it: those arrivals are referral, while the ranking the same placement supports produces organic ones. The same piece of work arrives under two labels, and a practice reading only their sum cannot separate the two effects.

Fields

ColumnTypeAliasDescription
channel_idSTRINGChannel IDPK. The identifier of this way of arriving.
channel_groupSTRINGChannel GroupThe kind of arrival: organic, referral, direct or ai.
sourceSTRINGSourceWho sent the visitor — google for organic search from Google, the domain for a referral.

Content Data Mart

Content

The text, images and video a site publishes: the first of the three pillars of search — what a business says about itself, and the material that gets indexed and served in AI answers and in Google. It is also one of the three things an SEO practice pays for, alongside the developer work that keeps the site in order and the placements bought on other people's sites, so this is where the first of those three cost lines sits.

Reading content as a volume is the reading this model is built to prevent. What gives a piece of content a chance of ranking is not that it exists, and not that a person rather than a machine produced it, but whether it adds knowledge to its topic. Content that adds nothing still costs what it cost to produce; and a large body of generated content that does not rank demonstrably worsens a site's traffic figures, often looking like a collapse.

Fields

ColumnTypeAliasDescription
content_idSTRINGContent IDPK. The identifier of this piece of content.
page_idSTRINGPage IDThe page this text, image or video sits on, and the only route from content to the site that publishes it. FK to Page
knowledge_gain_idSTRINGKnowledge Gain IDThe assessment of what new knowledge this content brings into its topic. FK to Knowledge Gain
content_typeSTRINGContent TypeWhat kind of content it is: text, image or video.
published_atTIMESTAMPPublished AtWhen this content went live on the page.
is_ai_generatedBOOLEANIs AI GeneratedWhether the content was produced by AI. On its own this decides nothing: what matters is whether the content adds knowledge to its topic.
production_costNUMERICProduction CostWhat it cost to create this content — one of the three things an SEO practice pays for, alongside developer work on the site and placements on other sites.

Relationships

Related data martOnCardinalityMeaning
Pagepage_id = page_idN:1The page this text, image or video sits on.
Knowledge Gainknowledge_gain_id = knowledge_gain_idN:1The assessment of what new knowledge this content adds.

Context Data Mart

Context

The topic a search belongs to — the subject someone is asking about, rather than the words they used to ask. Many different queries and many different keywords can be the same context: "web analytics for websites", "visit analytics for e-commerce" and "how do I find out where my customers came from" are all one topic, and a practice works on the topic rather than on the phrasings one at a time.

It is a mart of its own, with almost nothing on it, because it is what the rest of the model is judged against. Authority is contextual — standing is topical, and a site can carry it in one subject and none in another — and whether a page adds new knowledge is measured against a topic. Leaving the context out is what an entry-level practitioner does: calling a site authoritative because it is large and gets traffic, and buying a placement on it whose value in Google's eyes turns out to be quite moderate. The more often a site is cited within a context, the better its chance of coming up, and coming up high, for it.

Fields

ColumnTypeAliasDescription
context_idSTRINGContext IDPK. The identifier of this topic.
topicSTRINGTopicThe subject itself, in the practice's own words — the thing a set of queries, a set of keywords and a body of content are all about.

Crawl Data Mart

Crawl

One fetch of one page by one robot. A row is a single arrival, not a robot: a robot can fetch the same page more than once, and each fetch is its own row. Keeping the grain at the fetch is what lets crawling be read as behaviour over time rather than as a list of who is out there.

A crawler landing on a site learns its content and its links. The internal ones lead it to internal pages, which it also visits — each of those visits another row here. An external one may send it to a page it did not know, and that page belongs to somebody else and has no row in this model, so the fetch names a host and no page.

That is why this mart reaches the crawl rules by two routes rather than one. A fetch of the site's own page reaches them through the page and its site; a fetch of somebody else's page has only the route through the indexing request — one is the rules of the site whose page was fetched, the other the rules the run itself read.

Fields

ColumnTypeAliasDescription
crawl_idSTRINGCrawl IDPK. The identifier of this single fetch.
page_idSTRINGPage IDThe page that was fetched, when the fetch was of this site. Empty when the crawler followed a link off it onto somebody else's page. FK to Page
third_party_website_idSTRINGThird Party Website IDThe external host the robot went to, when the page fetched belonged to somebody else. FK to Third Party Websites
initiate_index_idSTRINGIndexing Request IDThe run that caused this fetch, which is how a request is tied to the pages it actually reached. FK to Initiate Index
crawler_user_agentSTRINGCrawler User AgentWho fetched the page, as it introduced itself. Anyone requesting a page on the internet is obliged to state a user agent, and usually that is more than enough to tell a robot from a person.
crawled_atTIMESTAMPCrawled AtWhen this fetch happened.
was_allowedBOOLEANWas AllowedThe crawl rules verdict on this fetch: whether the robot was let into the path it asked for.
links_foundINTEGERLinks FoundHow many links the crawler came away with. Internal ones lead it to pages it then visits too; external ones either credit a page it already knew or send it to crawl one it did not.

Relationships

Related data martOnCardinalityMeaning
Pagepage_id = page_idN:1The page fetched, when the fetch was of this site; empty for a third-party page.
Third Party Websitesthird_party_website_id = third_party_website_idN:1The external site fetched, when the crawler followed a link off this one.
Initiate Indexinitiate_index_id = initiate_index_idN:1The indexing run this fetch belongs to.

Google Search Console Data Mart

Google Search Console

The property registered with Google: Google's own facility where a site is entered directly, as a claim of ownership and a request that it be crawled — and Google goes and crawls it. It is the one place in this model where a site is declared to a search engine instead of being discovered by one, which is why it is held as an object of its own rather than as a flag on the site.

It carries no numbers. Impressions, clicks and positions are read through this property and are kept in Search Performance; what lives here is the registration itself — which site, whether the claim was confirmed, and when. Search Console is also where searching done by AI is fairly visible, as very long queries a normal person would never write.

Fields

ColumnTypeAliasDescription
property_idSTRINGProperty IDPK. The identifier of this registered property.
website_idSTRINGWebsite IDThe site this property was registered for. FK to Website
property_urlSTRINGProperty URLThe address the property covers, as it was entered when the site was registered.
is_verifiedBOOLEANIs VerifiedWhether ownership of the site has been confirmed, which is what makes the property a claim rather than an entry.
verified_atTIMESTAMPVerified AtWhen ownership was confirmed, so a change in crawling or in search numbers can be lined up against the moment the site was declared.

Relationships

Related data martOnCardinalityMeaning
Websitewebsite_id = website_idN:1The site this property was verified for.

Initiate Index Data Mart

Initiate Index

The moment a site or one of its pages enters the queue to be indexed. The name is the one coined on the call this model was drawn from rather than a phrase used elsewhere in the trade.

A run arrives here by one of three routes. A site can be entered into Google's Search Console directly, as a claim of ownership and a request that it be crawled. A sitemap can offer the structured list of pages to be indexed. Or a crawler already working somewhere else can find a link to a page it did not know, and go and crawl that page. The sitemap route is optional: a sitemap helps, but if there is none, Google will crawl the site anyway, as long as the robots.txt is in order.

The date here is when the asking happened, not when a robot turned up, and the gap to the fetches that followed is what shows how long a site waited.

Fields

ColumnTypeAliasDescription
initiate_index_idSTRINGIndexing Request IDPK. The identifier of this run.
robots_txt_idSTRINGRobots File IDThe crawl rules the arriving crawler reads on the way in, whichever route asked for the run. FK to robots.txt
requested_atTIMESTAMPRequested AtWhen the site or page was put into the queue, which is not when a robot arrived.
triggerSTRINGTriggerWhich route started the run: search_console for a request made directly in the registered property, sitemap for the page list being offered, external_link for a crawler that met a link to a page it did not know.

Relationships

Related data martOnCardinalityMeaning
robots.txtrobots_txt_id = robots_txt_idN:1The rules the run read before fetching anything.

Keyword Data Mart

Keyword

A term the practice targets — the middle of the three levels this model keeps for demand. It is one of the main terms of the trade. A keyword is, to a considerable extent, a synonym of a topic: many keywords can be one and the same context, and the practice works on the topic rather than on each phrasing separately.

The word carries two meanings in everyday use, and this model splits them. Asked directly whether the thing a user sees is a keyword or a search query, the answer on the record was: it is a search query. So the string somebody typed is a Search Query, and keyword is used here by this model's own convention for the term the practice targets — the entry on the practice's list, not the entry in somebody's search box. Holding the two apart is what makes it possible to ask which of the chosen terms nobody searches, and which searches a site picks up that it never aimed at.

Fields

ColumnTypeAliasDescription
keyword_idSTRINGKeyword IDPK. The identifier of this term.
context_idSTRINGContext IDThe topic this term belongs to; a keyword is to a considerable extent a synonym of a context, and many keywords can be one and the same one. FK to Context
keyword_textSTRINGKeywordThe term itself, as the practice writes it on its own list.
is_targetedBOOLEANIs TargetedWhether the practice is optimising for this term, as opposed to keeping it on the list to watch.

Relationships

Related data martOnCardinalityMeaning
Contextcontext_id = context_idN:1The topic this term belongs to; many keywords share one.

Knowledge Gain Data Mart

Knowledge Gain

Whether a piece of content brings new knowledge into a topic — a super-important component of ranking. Google holds patented technology for the judgement, published information rather than a secret, by which it assesses how valuable a page is in bringing new knowledge into a context. It is kept as an object of its own rather than as a column on the content because there is a context, there is content other sites have already written on that context, and what is being measured is the difference a page makes to it.

It is the brick people fail to understand when they hand content creation over to AI and are then surprised at how their site is doing. The term is not much circulated, and it stays uncirculated because it gets in the way of selling AI content generators. It is not a verdict on the author: content that merely retells what is already available carries no gain, and content built on data nobody else holds carries it whether or not a machine wrote the words.

Fields

ColumnTypeAliasDescription
knowledge_gain_idSTRINGKnowledge Gain IDPK. The identifier of this assessment.
context_idSTRINGContext IDThe topic the content is being judged against — there is no such thing as a gain in the abstract. FK to Context
scoreFLOATKnowledge Gain ScoreHow much new knowledge this content brings into its topic, as a number. It is read against the topic, not against the content on its own.
adds_new_knowledgeBOOLEANAdds New KnowledgeWhether the content brings knowledge that was not already in general circulation, or only retells what other sites have written on the topic.
is_expert_authoredBOOLEANIs Expert AuthoredWhether the author is already visible in this topic as the author of content, ideas, opinions and analysis — which gives a page a chance of ranking despite weak support from links.
is_based_on_unique_dataBOOLEANIs Based on Unique DataWhether the content rests on data that was not otherwise available. Where it does, machine authorship makes no difference to how it ranks.
contradicts_consensusBOOLEANContradicts ConsensusWhether the content argues against an established consensus in its topic. Content that does is deranked even when it is right.

Relationships

Related data martOnCardinalityMeaning
Contextcontext_id = context_idN:1The topic the gain is measured against; the same content scores differently in another.

Page Data Mart

Page

One page of the site being optimised: the unit that gets crawled, indexed, ranked and arrived at. A page is what content sits on, and the level at which what a site asks a search engine for — index this — can be compared with what happened to it. is_in_sitemap is a column here and not a link to the page list: the list is optional, and being named in it is simply true or false about a page.

It carries its own behaviour numbers because Google reads user signals per page and for the site as a whole, as an average. A hack turns on that arithmetic: removing content that does not work lifts the averages of what is left, which Google reads as an improvement in quality. Whether a dead page should be removed, redirected or rewritten is not settled, and depends heavily on how spoilt Google is for content on that topic: where it has plenty, such pages should not be on the site at all; where it has little, it will show even weak ones for want of anything better. Pages kept for regulatory or documentation reasons can be held quite safely.

Fields

ColumnTypeAliasDescription
page_idSTRINGPage IDPK. The identifier of this page.
website_idSTRINGWebsite IDThe site this page belongs to, which is what separates the pages of the site being optimised from anyone else's. FK to Website
urlSTRINGPage URLWhere the page is published.
page_typeSTRINGPage TypeWhat kind of page it is: home, service, contact, case_study or blog.
is_indexableBOOLEANIs IndexableWhether the page is telling search engines it wants to be indexed.
is_indexedBOOLEANIs IndexedWhether the page is actually in the index. It can differ from what the page asks for.
is_in_sitemapBOOLEANIs In SitemapWhether the site's page list names this page. The list itself is optional — a site with none is still crawled.
last_crawled_atTIMESTAMPLast Crawled AtWhen a search robot last fetched this page.
avg_time_on_pageFLOATAverage Time on PageHow long visitors stay on this page — a user signal Google reads per page as well as averaged over the site.
avg_depth_of_scrollFLOATAverage Depth of ScrollHow far down this page visitors get before they stop.
bounce_rateFLOATBounce RateThe share of visitors who arrive on this page and leave straight away.
has_organic_trafficBOOLEANHas Organic TrafficWhether the page earns any traffic from search at all. Pages that earn none are the ones a removal experiment is aimed at, and the ones that move the site's averages down.

Relationships

Related data martOnCardinalityMeaning
Websitewebsite_id = website_idN:1The site this page belongs to, which is what separates the pages of the site being optimised from anyone else's.

robots.txt Data Mart

robots.txt

The file a site publishes to tell arriving search robots where they may go and where they may not, and which of them the owner wants indexed at all. It is its own object rather than a field on the site because it decides what everything downstream is even allowed to see.

A site that publishes no sitemap is still crawled, as long as its robots.txt is in order. A path closed here is a path the crawler does not fetch, whatever the page behind it says about wanting to be indexed — which is why the rules and the pages are worth reading together rather than separately. The count of closed paths says nothing on its own about whether the closures are right: a site that deliberately keeps part of itself out of the index has a high count and a sound setup, and a site that has closed a content directory by mistake looks identical from here.

Fields

ColumnTypeAliasDescription
robots_txt_idSTRINGRobots File IDPK. The identifier of this crawl-rules file.
urlSTRINGFile URLWhere the file is published, at the root of the site it governs.
is_validBOOLEANIs ValidWhether the file parses as valid crawl rules. When it does not, an arriving crawler is left without the instructions it came for.
disallowed_path_countINTEGERDisallowed PathsHow many paths the file closes to crawlers. Read against what the pages behind them are meant to do, not on its own.
last_modified_atTIMESTAMPLast Modified AtWhen the rules last changed, so a shift in crawl behaviour can be lined up against a change in what the crawler was told.

Search Performance Data Mart

Search Performance

Where one page stood for one query on one day, and what that standing earned. A page can hold quite different positions for two queries on the same day, so the key is all three.

The distinction this model turns on lives here: a position is not traffic. Coming from nowhere into the top 20 is a very noticeable result, and there is no traffic yet. Getting from the top 20 into the top five is another three or four months of work: 90% of traffic is concentrated in Google's first five lines, because people do not scroll down when the first five answer them. Searches made by AI land here beside the ones people typed, and a position averaged over both puts the two kinds into one figure.

Because every row is dated, the two clocks of SEO can be told apart here. Work done fundamentally — strong content, good links, and no competitor whose own work starts pushing a site off the top places — can stay stable for as much as five years. A hack can be closed off by an algorithm update, after which it stops giving what it was giving.

Fields

ColumnTypeAliasDescription
dateDATEDatePK. The day these figures were measured for, which is what makes a position readable as movement rather than as a single snapshot.
search_query_idSTRINGSearch Query IDPK. The query this row's impressions and position were measured for. FK to Search Query
page_idSTRINGPage IDPK. The page that appeared in the results for that query. FK to Page
impressionsINTEGERImpressionsHow many times the page was shown in the results for this query on this day.
clicksINTEGERClicksHow many times someone went from the results to the page.
ctrFLOATClick-Through RateThe share of impressions that became clicks.
avg_positionFLOATAverage PositionWhere the page stood in the results that day, averaged over its impressions. It is not traffic: a page in the top 20 has a real result and no visits yet, while 90% of traffic sits in the first five lines.

Relationships

Related data martOnCardinalityMeaning
Search Querysearch_query_id = search_query_idN:1The query this row's impressions and position were measured for.
Pagepage_id = page_idN:1The page that appeared in the results for it.

Search Query Data Mart

Search Query

What somebody actually put into a search box. It is the finest of the three levels this model keeps for demand: a query is the string that was typed, a keyword is — by this model's own convention — the term the practice targets, and a context is the topic both of them sit in.

The level is where ranking is settled: a position is a position for a query. In the practitioner's own picture of it, for a given query a site was not taking part in the lottery at all, and now it is in it and holding a place.

The boundary between a query and a keyword has blurred a little. A search made by AI is, as a rule, not a keyword but a pile of key words — a very long string a normal person would never write — and those searches land in this mart next to the short ones a person typed.

Fields

ColumnTypeAliasDescription
search_query_idSTRINGSearch Query IDPK. The identifier of this query string.
keyword_idSTRINGKeyword IDThe term the practice targets that this typed query was matched to. FK to Keyword
query_textSTRINGSearch QueryThe string that was actually searched for.
is_ai_generated_queryBOOLEANIs AI-Generated QueryWhether the query looks like a search made by AI rather than typed by a person: as a rule not a keyword but a pile of key words, a very long string a normal person would never write.

Relationships

Related data martOnCardinalityMeaning
Keywordkeyword_id = keyword_idN:1The term the practice targets that this typed query was matched to.

sitemap.xml Data Mart

sitemap.xml

The structured list of its own pages a site hands to search engines for indexing. It is the site saying, in one place, what it consists of — which is help a crawler can use, but not help it depends on: removing the sitemap does not worsen a site's results in search, as long as everything else is in order.

A site may publish none at all, and the absence is a legitimate state rather than a gap in the data, which is why the list is held separately rather than folded into the site record. Where a sitemap does exist, the number it carries is a claim about what the site contains, and a claim that has drifted away from the pages actually published is worth knowing about. Publishing content and telling search engines about it are two separate acts, and the gap between them is visible here.

Fields

ColumnTypeAliasDescription
sitemap_idSTRINGSitemap IDPK. The identifier of this page list.
urlSTRINGSitemap URLWhere the list is published for search engines to fetch.
listed_page_countINTEGERListed PagesHow many pages the list offers for indexing — the site's own claim about what it consists of, to be read against the pages it actually publishes.
last_submitted_atTIMESTAMPLast Submitted AtWhen the list was last handed to search engines, which is a separate act from publishing the pages in it.

Third Party Websites Data Mart

Third Party Websites

The sites a practice does not own: the hosts that link to it, mention it, or could. Off-site work is the third pillar of search — what the rest of the web says about a business.

It deliberately carries no single standing for the site: standing combines the organic traffic a host already has with the links pointing at it within a topic, so it is measured per site and per topic and kept on its own object. What lives here is what is true of the host whatever the topic — its traffic, and whether it is one of the broadly authoritative sites, Forbes and Wikipedia, where being mentioned in any context counts heavily in a site's favour.

Two things have made this work much more complicated than it used to be. Buying links by the hundred used to be the whole game, and that is no longer how it works; reputation and topical relevance are what count now. And AI does not crawl sites, at least not yet: what it responds to is a brand being mentioned in a context, and, judging by indirect signals, the reputation of the source matters a great deal.

Fields

ColumnTypeAliasDescription
third_party_website_idSTRINGThird Party Website IDPK. The identifier of this external host.
domainSTRINGDomainThe host itself, as it appears on the placements that sit on it.
organic_trafficINTEGEROrganic TrafficThe search traffic the host already receives — one of the two ingredients behind its standing, the other being the links it holds within a topic.
is_general_authorityBOOLEANIs General AuthorityWhether this is one of the broadly authoritative hosts that count in any context, rather than in one topic.

Visit Data Mart

Visit

Somebody — or something — arriving on one of a site's pages. This is the event the whole model points at: most practitioners do treat the visit as their target event, and it is where this model stops. What happens past the arrival belongs to a different model.

There is no key from a visit to the query that produced it, and that is deliberate rather than missing. What an arrival brings with it is how it arrived — the channel and the source — and not the words somebody typed to get there. Demand and behaviour therefore meet in the daily aggregate in Search Performance, which is keyed by query and page, and nowhere else.

A visit is also the raw material of the user signals that decide whether a position survives: of the sites somebody opens for one query, the one they spent three minutes on and the one they left at once are read differently, and the one they left at once is thrown out of the top five so the next site can have its chance.

Fields

ColumnTypeAliasDescription
visit_idSTRINGVisit IDPK. The identifier of this arrival.
page_idSTRINGPage IDThe page that was landed on. FK to Page
channel_idSTRINGChannel IDHow the visitor arrived, as an analytics tool classifies it. FK to Channel
occurred_atTIMESTAMPOccurred AtWhen the arrival happened.
user_agentSTRINGUser AgentHow the visitor introduced itself. Anyone requesting a page on the internet is obliged to state a user agent.
is_botBOOLEANIs BotWhether this was a robot rather than a person, read off the user agent — usually more than enough to tell the two apart.
pages_viewedINTEGERPages ViewedHow many pages the visitor opened on the site during this arrival.
duration_secondsINTEGERDuration in SecondsHow long the arrival lasted.
bouncedBOOLEANBouncedWhether the visitor left straight away instead of going any further into the site.
is_qualifiedBOOLEANIs QualifiedWhether this arrival met the practice's own definition of a visit worth having. It is the last thing the model says about a visit.

Relationships

Related data martOnCardinalityMeaning
Pagepage_id = page_idN:1The page the visit landed on.
Channelchannel_id = channel_idN:1How the visitor arrived.

Website Data Mart

Website

The site being optimised — the object the whole technical side of search work is done to, and the boundary between what is within a practice's control and what is not. Content is placed on it, and its crawl rules and page list belong to it; the part of SEO that concerns a site's own condition ends where this record ends.

Google sees how people behaved after arriving, and it looks at those user signals both per page and for the site overall, as an average — so speed, layout and the behaviour numbers are properties of the site rather than of any one page. A site need not be superfast, but it does have to be reasonably fast, and 80% of its traffic may be coming from phones while its layout breaks on them. The two connect: a visitor who opens the site on a phone and leaves at once, because the content they came for could not be read, earns the site a minus for user signals, and the site slides down in search.

Fields

ColumnTypeAliasDescription
website_idSTRINGWebsite IDPK. The identifier of this site.
domainSTRINGDomainThe domain the site is published under — the name everything else in the model is claimed and measured against.
robots_txt_idSTRINGRobots File IDThe crawl-rules file published at the root of this site, which tells an arriving robot where it may go. FK to robots.txt
sitemap_idSTRINGSitemap IDThe list of its own pages this site offers for indexing. Empty when the site publishes none, which does not stop the site being crawled, as long as its robots.txt is in order. FK to sitemap.xml
page_speed_scoreFLOATPage Speed ScoreHow quickly the site loads. The bar is not superfast — it is reasonably fast.
mobile_optimization_scoreFLOATMobile Optimization ScoreHow well the layout holds up across device types, read against how much of the site's traffic arrives from phones.
avg_time_on_pageFLOATAverage Time on PageTime on page averaged across the site — one of the user signals Google reads for the site as a whole and not only per page.
avg_depth_of_scrollFLOATAverage Depth of ScrollHow far down a page visitors get, averaged across the site.
bounce_rateFLOATBounce RateThe share of arrivals that leave straight away, whatever sent them away — being unable to read the content on a phone is one such reason.
development_costNUMERICDevelopment CostWhat has been spent on developer work to put the site in technical order and keep it there.

Relationships

Related data martOnCardinalityMeaning
robots.txtrobots_txt_id = robots_txt_idN:1The file that tells crawlers where on this site they may go.
sitemap.xmlsitemap_id = sitemap_idN:1The optional index of the site's pages; empty when the site publishes none.

Apply to your project

  1. 1

    Install the Import Model plugin

    One plugin, installed once, in your own OWOX workspace.

    Get the plugin →

  2. 2

    Import this model

    Point it at this bundle and it creates every data mart above, joins and all.

    Open the model →

  3. 3

    Plug in your data and destinations

    Connect your own sources and send the results where your team already works.

    Browse connectors →

4. Optional — customize as you wish. Rename a column, drop a mart, add your own: once it is imported it is yours, and nothing here syncs back.