Original research by Kyle Alm · 2026-08-16
Original research: how spamdexing became SEO, dated to the hour
Wikipedia does not coin language; it admits it. New ideas do not appear in a reference work — established ones do — so every arrival and departure in its text is a dated acceptance ruling. We measured 5,901,821 words of revision history to chart how the industry’s phrases rose, crossed, and died.
What this research found
On 24 February 2003, at 22:58 UTC, an editor using the account Sherwoodseo created the Wikipedia article Search engine optimization — that link opens the exact revision this study measured, not the article as it stands today.
The account name is the point. SEO was already the word. Someone had it in their username before the encyclopedia had an article for it, and the article they wrote describes an established consulting industry with firms, clients and rival techniques. Nothing was invented that night. What happened was narrower and more interesting: a term already in professional use crossed the threshold into the reference work that ratifies English.
The encyclopedia had described the practice long before it accepted the profession. The article Spamdexing was created on 10 March 2002 — eleven and a half months earlier — and its first revision explains modifying pages to rank higher, invisible keyword text, and search engines removing offenders. It never once uses the words “optimization” or “SEO”. Wikipedia had a word for gaming search engines a year before it had a word for the job.
Part one
Naming decisions record what a term meant. Usage records whether anyone still says it. Measuring each phrase at density per 1,000 words — never raw counts, since the corpus grows fivefold — gives every piece of trade vocabulary a full 22-year curve. The clearest reading is peak year, and how much of that peak survives today.
| Year | founding era | middle era | modern era |
|---|---|---|---|
| 2004 | 3.00 | 0.03 | 0.03 |
| 2005 | 3.28 | 0.21 | 0.08 |
| 2006 | 2.81 | 0.30 | 0.36 |
| 2007 | 2.81 | 0.56 | 0.48 |
| 2008 | 2.61 | 0.67 | 0.65 |
| 2009 | 2.39 | 0.73 | 0.77 |
| 2010 | 2.45 | 0.80 | 0.78 |
| 2011 | 2.32 | 0.66 | 1.10 |
| 2012 | 2.30 | 0.73 | 1.37 |
| 2013 | 2.13 | 0.57 | 1.51 |
| 2014 | 2.05 | 0.55 | 1.91 |
| 2015 | 2.01 | 0.51 | 2.30 |
| 2016 | 1.89 | 0.48 | 2.81 |
| 2017 | 1.82 | 0.46 | 3.00 |
| 2018 | 1.77 | 0.42 | 2.74 |
| 2019 | 1.76 | 0.41 | 2.79 |
| 2020 | 1.77 | 0.45 | 3.05 |
| 2021 | 1.70 | 0.42 | 3.21 |
| 2022 | 1.71 | 0.42 | 3.28 |
| 2023 | 1.62 | 0.38 | 3.32 |
| 2024 | 1.66 | 0.37 | 2.99 |
| 2025 | 1.63 | 0.37 | 3.00 |
| 2026 | 1.63 | 0.35 | 2.87 |
The drift has a direction, and it is not a rebrand. The vocabulary that entered with the field described what a practitioner does to a page: meta keywords, keyword density, title tags, cloaking, doorway pages. The vocabulary of the 2020s describes what a system does: machine learning, artificial intelligence, large language models, structured data. The subject of the sentence changed.
Faded — below 35% of peak · 10 terms
Enduring — at or above 35% · 23 terms
Blink-and-gone — too few observations to draw
One tile per tracked term, 2004–2026; dot marks the peak year, percentage is how much of that peak survives in 2026. Each tile is scaled to its own peak — read the shapes and dates, not the heights; the hover text carries each term’s absolute density. A term seen in fewer than four years gets a chip, not a curve — a line through three points invites a trend the data cannot support.
The phrase “keyword density” peaked in 2010 and has lost about four-fifths of its frequency since — the single most-asked question about this dataset, and the answer is a date. But the pattern above it is sharper. Meta keywords, cloaking, doorway pages and page rank all peak in 2004–2006, the corpus’s earliest years, and never recover. The vocabulary that arrived with the field is the vocabulary that died.
Two spelling cases are worth separating. pagerank as one word is the largest early term in the whole set, peaking at 1.387 per 1,000 words in 2006, and retains 46%. page rank as two words retains 10%. The proper noun survived; spelling it out as a generic technique did not. Likewise serps retains 36% while serp singular retains 51% — the plural was jargon, the singular became a noun.
Social media is the largest term in the entire set by an order of magnitude, peaking at 2.684 per 1,000 words. And the terms peaking in 2025 and 2026 are the live vocabulary: artificial intelligence, machine learning, user experience, structured data, large language. Two of them — large language and core web vitals — do not appear anywhere in the corpus before 2022.
Below a third of peak a phrase has effectively left the working vocabulary, even though it still appears in historical passages. That 35% line is a presentation choice, stated so you can move it: the full ordering runs continuously from 5% to 100%.
Every term below is defined in our SEO glossary, which carries all 36 definitions alongside these dates.
| Term | First | Peak | Peak /1k | 2026 /1k | Retained | |
|---|---|---|---|---|---|---|
| link juice | 2015 | 2015 | 0.004 | 0.000 | 0% | |
| alt text | 2021 | 2021 | 0.003 | 0.000 | 0% | |
| invisible text | 2005 | 2005 | 0.026 | 0.000 | 0% | |
| meta keywords | 2006 | 2006 | 0.116 | 0.005 | 5% | |
| cloaking | 2004 | 2004 | 0.256 | 0.024 | 9% | |
| doorway pages | 2004 | 2004 | 0.051 | 0.005 | 10% | |
| page rank | 2004 | 2004 | 0.128 | 0.013 | 10% | |
| meta tags | 2004 | 2004 | 0.384 | 0.075 | 19% | |
| keyword stuffing | 2005 | 2005 | 0.091 | 0.019 | 21% | |
| keyword density | 2005 | 2010 | 0.115 | 0.024 | 21% | |
| link farm | 2004 | 2007 | 0.120 | 0.027 | 22% | |
| title tag | 2005 | 2006 | 0.044 | 0.011 | 24% | |
| white hat | 2004 | 2006 | 0.124 | 0.043 | 34% |
| Term | First seen | Peak | Peak /1k | Retained | |
|---|---|---|---|---|---|
| structured data | 2004 | 2026 | 0.053 | 100% | |
| core web vitals | 2022 | 2024 | 0.008 | 99% | |
| large language | 2022 | 2025 | 0.128 | 98% | |
| artificial intelligence | 2006 | 2025 | 0.186 | 94% | |
| user experience | 2005 | 2025 | 0.139 | 92% | |
| machine learning | 2005 | 2025 | 0.093 | 91% | |
| content marketing | 2008 | 2020 | 0.202 | 91% | |
| search engine results | 2004 | 2007 | 0.233 | 86% | |
| mobile-first | 2016 | 2025 | 0.013 | 80% | |
| link building | 2005 | 2020 | 0.074 | 79% | |
| social media | 2006 | 2022 | 2.684 | 76% | |
| black hat | 2004 | 2017 | 0.106 | 76% | |
| paid links | 2007 | 2007 | 0.045 | 59% | |
| paid inclusion | 2004 | 2005 | 0.208 | 56% | |
| serp | 2006 | 2012 | 0.193 | 51% | |
| anchor text | 2004 | 2005 | 0.208 | 50% | |
| search engine ranking | 2005 | 2007 | 0.045 | 47% | |
| pagerank | 2004 | 2006 | 1.387 | 46% | |
| organic search | 2005 | 2005 | 0.221 | 46% | |
| backlinks | 2004 | 2005 | 0.247 | 43% | |
| serps | 2004 | 2005 | 0.234 | 36% | |
| reciprocal links | 2005 | 2007 | 0.015 | 35% | |
| nofollow | 2005 | 2010 | 0.431 | 35% |
Part two
Trade language rarely just dies — it gets replaced. Tracking the old form and the new form of the same idea through the same corpus puts a date on the moment the industry’s language actually turned over.
The hyphen died in 2011. “email marketing” overtook “e-mail marketing” that year and never fell behind again — a spelling reform you can date to the year. The bigger renaming took longer: digital marketing overtook internet marketing in 2016, the field’s older name for itself giving way in place.
The other two panels show the limit of the record, honestly drawn. website already led web site when the corpus opens in 2004, and web crawler already led search engine spiders — those races were decided before our first snapshot, so the chart can only say “already over,” not when.
Wikipedia does not coin language; it admits it. A phrase’s first appearance in this corpus is an acceptance date — the year the encyclopedia judged the term established enough to use — and lining those dates up shows acceptance arriving in waves, not a trickle.
The striking feature is the gap: nothing new got in for seven years. Of the 27 still-current phrases on the timeline, 23 arrive by 2008 — then from 2009 through 2015, not one new phrase enters the corpus. Admission resumes with mobile-first (2016) and general data protection regulation (2017), and the AI-era vocabulary lands as a wave in 2022. The freeze is a claim about this study’s vocabulary — the tracked list plus the discovered risers — not about every phrase in Wikipedia, but it matches the acceptance thesis exactly: an encyclopedia admits a field’s language in bursts, when the language has already settled elsewhere.
The list above was chosen by hand. This one was not: every two- and three-word phrase in the corpus was compared between 2004–2009 and 2021–2026, and these are the ones that fell hardest — each appearing in at least eight articles early on, so no single editor’s habit can masquerade as industry-wide jargon. Read together, they show that trade language dies three distinct ways.
| Phrase | Articles | Early /1k | Late /1k | Retained | |
|---|---|---|---|---|---|
| e-mail spam | 11 | 0.117 | 0.011 | 10% | |
| e-mail marketing | 11 | 0.157 | 0.017 | 11% | |
| robots exclusion standard | 8 | 0.094 | 0.017 | 18% | |
| google yahoo msn | 15 | 0.069 | 0.016 | 24% | |
| yahoo search marketing | 12 | 0.078 | 0.019 | 24% | |
| open directory project | 12 | 0.066 | 0.016 | 25% | |
| instant messaging | 13 | 0.094 | 0.025 | 27% | |
| live search | 13 | 0.310 | 0.086 | 28% | |
| search engine spiders | 9 | 0.069 | 0.021 | 31% | |
| world wide web | 47 | 1.075 | 0.349 | 32% | |
| ask jeeves | 10 | 0.123 | 0.041 | 34% | |
| click through | 11 | 0.066 | 0.022 | 34% | |
| internet marketing | 41 | 0.364 | 0.125 | 34% | |
| markup language | 8 | 0.159 | 0.058 | 36% | |
| image search | 12 | 0.063 | 0.024 | 38% |
Compression. Robots exclusion standard keeps under a fifth of its old frequency; world wide web and markup languageroughly a third. Nobody writes those out any more — they write robots.txt, the web, HTML. The concept survives; the words collapse into an abbreviation.
Product death. Yahoo search marketing, open directory project, ask jeeves, live search — the jargon died with the products it named. And the three-name reflex of the era, google yahoo msn, is nearly gone: naming a search engine used to mean naming three, and now it means naming one.
The trade’s own words, fading in place. Internet marketing — the field’s older name for itself — keeps a third of its former use. And search engine spiders, the exact phrase Sherwoodseo created as a redirect at 23:45 on the founding night, keeps under a third. The jargon he wired into the encyclopedia is now fading out of it.
And sometimes the trade’s word wins. The most formal name in the early corpus, unsolicited commercial e-mail, keeps 2% of its 2004 frequency — the encyclopedia now just writes spam, the word the trade always used. Hyper text markup language spelled out is gone entirely, last seen 2019. Compression wins even against formality.
Part three
Sherwoodseo made 102 edits in their entire Wikipedia lifetime. Nine of them, in sixty minutes, moved the trade’s vocabulary into the encyclopedia.
Four of the nine edits are not about SEO as a topic at all. They reach into articles the encyclopedia already had — Web crawler, Meta element, Search engine — and connect the trade’s vocabulary to them. The two redirects work in that direction: a new page called Search engine spider was created whose entire content pointed readers to the existing Web crawler article. The practitioner’s word for the thing was filed as another name for the encyclopedia’s word.
This was reviewed. At 23:56 the new editor left a one-word message on their own talk page; within the hour an established editor, Hephaestos, replied “Hello! Welcome to the ‘pedia”. The next morning Hephaestos edited both Meta element and Search engine (computing) — the two existing articles Sherwoodseo had touched. The additions stood. That is what acceptance looks like in a reference work: a newcomer’s contribution inspected by a regular and left in place.
The first revision opens by dating the practice to the mid-1990s and carries no citations at all. An unsourced origin story, written by a practitioner, became the encyclopedia’s account of when the field began — and is still roughly the story told today.
Part four
Once a term is in, the encyclopedia has to rule on its rivals. The SEO article has accumulated 79 dated redirects, and read chronologically they are a six-year sequence of verdicts — each one an editor deciding that some other name for the work meant this one.
| Date | Term absorbed | Decided by |
|---|---|---|
| 2004-05-17 | Searchability optimization | Jmabel |
| 2005-08-12 | Search engine engineering | 66.120.119.115 |
| 2005-12-03 | Search engine ranking | Search Engines Web |
| 2006-06-01 | Search Engine Placement | Guruweb |
| 2006-06-01 | Search Engine Position | Guruweb |
| 2006-11-19 | Search engine prominence | Sepguy |
| 2007-04-14 | Search Optimization Marketing | RHaworth |
| 2010-03-17 | Key search engineering | Dannycarney |
| 2010-10-31 | Digital Asset Optimization | Cesareone |
Key search engineering is the most revealing of them. It was not a convenience redirect — it was an 8,646-character article whose opening sentence is a near-verbatim fork of the SEO article’s lead. In March 2010 somebody tried to rename the entire field and was absorbed back within the year.
The acronym took far longer to be accepted than the phrase. Seo pointed at the article from day one in 2003, but SEO in capitals was not pointed there until 13 March 2011 — and it was contested, reverted to the disambiguation page three weeks later, then re-pointed. Eight years after the industry was using the acronym daily, Wikipedia was still arguing about whether it primarily meant this.
Part five
A citation is a link out of the encyclopedia. Tracked over 25 years, those links tell a second story about search-marketing sources: most of them do not last. The same corpus that retired the founding vocabulary also retired the founding citations — language and sources turn over together.
More than half of every domain the encyclopedia ever cited on these topics has since been dropped, and the typical source survives just 4 of 23 years. But the attrition is not uniform — it depends heavily on when a source first appeared.
Every cohort is tracked over the same eight-year window, so the comparison is fair. Of the domains first cited in 2006, fewer than half survived even a single year, and 30.1% lasted the eight. The 2018 cohort held 55.4% over the same span. Stretch the 2006 line all the way to today and only 25.3% remain. Early search-marketing citations were disposable in a way later ones are not — the blog posts, forum threads and defunct tools of the mid-2000s, cited once and replaced.
External citations nearly tripled — 13.6 to 36.6 per 1,000 words — while internal wikilinks fell by a third. In 2004 these articles linked inward 6.14 times for every outward citation; by 2026 that ratio is 1.55. The subject went from something Wikipedia explained mostly by cross-reference to something it documented from outside sources — the same shift, measured a third way.
| Year | External | Internal | Internal : external |
|---|---|---|---|
| 2004 | 13.6 | 83.6 | 6.14 |
| 2005 | 13 | 68.7 | 5.29 |
| 2006 | 15.4 | 72.6 | 4.72 |
| 2007 | 17.5 | 79.3 | 4.54 |
| 2008 | 18 | 75.7 | 4.21 |
| 2009 | 19.3 | 70.6 | 3.66 |
| 2010 | 21.1 | 72.4 | 3.42 |
| 2011 | 22.5 | 70.6 | 3.13 |
| 2012 | 23.7 | 71.7 | 3.03 |
| 2013 | 27.3 | 56.1 | 2.05 |
| 2014 | 27.7 | 54.6 | 1.97 |
| 2015 | 28.3 | 53.5 | 1.89 |
| 2016 | 27.4 | 50.8 | 1.85 |
| 2017 | 28.8 | 50.9 | 1.77 |
| 2018 | 29.5 | 51 | 1.73 |
| 2019 | 30.1 | 52.5 | 1.74 |
| 2020 | 31.2 | 53.1 | 1.7 |
| 2021 | 31.6 | 53.8 | 1.7 |
| 2022 | 32.8 | 54.1 | 1.65 |
| 2023 | 34.4 | 54.6 | 1.59 |
| 2024 | 35.1 | 56 | 1.6 |
| 2025 | 35.8 | 56.4 | 1.57 |
| 2026 | 36.6 | 56.8 | 1.55 |
Part six
A reference work shows what it accepts by what it cites. Start with the first bibliography the field ever had — and what became of it.
The 24 February 2003 revision carries zero external links. It was written from the author’s own knowledge, including the claim about when the practice began. The first seven citations appear in the 2004 snapshot, and they are the opening reading list for the whole discipline.
| Cited domain | Status today | What it serves now |
|---|---|---|
| searchenginewatch.com | — | |
| google.com | Redirects elsewhere | developers.google.com |
| realseo.com | Host placeholder page | Welcome realseo.com - BlueHost.com |
| threadwatch.org | Redirects elsewhere | internetmarketingninjas.com |
| seopages.com | Domain gone | DNS does not resolve |
| toprank.blogspot.com | Minneapolis SEO Blog | |
| seo-glossary.com | Resolves, serves nothing | — |
6 of the 7 domains still answer. Only 2 are still operating as themselves. That gap is the whole point of checking. A link checker counting HTTP 200s would report six survivors: realseo.com returns a hosting-provider placeholder, seo-glossary.com returns 163 empty bytes, and threadwatch.org redirects into the company that absorbed it. Only searchenginewatch.com — publishing same-day news when we checked — and toprank.blogspot.com are still doing what they were cited for.
The one genuinely dead entry is seopages.com: its name no longer resolves in DNS at all. And the most durable is the one nobody would call a citation — google.com/webmasters/seo.html still resolves, twenty-two years later, by redirecting to Google’s current SEO starter guide. The platform kept its content addressable across two decades of reorganisation; most of the independent trade did not.
One hypothesis this study set out to test was that large platforms have come to dominate these pages. The data does not support it, and points the other way.
| Year | Citations | Distinct domains | Top-10 share | HHI |
|---|---|---|---|---|
| 2004 | 532 | 363 | 16% | 0.0065 |
| 2010 | 3,871 | 1,578 | 18.1% | 0.0055 |
| 2016 | 8,246 | 2,955 | 13.3% | 0.0033 |
| 2020 | 10,969 | 3,356 | 15.1% | 0.0041 |
| 2026 | 13,726 | 3,685 | 15.8% | 0.0045 |
| Year | 1st | 2nd | 3rd | 4th |
|---|---|---|---|---|
| 2004 | google.com 4.9% | w3.org 2.3% | nytimes.com 1.7% | wired.com 1.3% |
| 2010 | google.com 3.6% | forbes.com 3.2% | nytimes.com 2.2% | w3.org 2.1% |
| 2016 | google.com 2% | nytimes.com 1.9% | techcrunch.com 1.5% | w3.org 1.4% |
| 2020 | nytimes.com 2.8% | techcrunch.com 1.7% | google.com 1.5% | theverge.com 1.5% |
| 2026 | nytimes.com 3.1% | theverge.com 2% | techcrunch.com 1.7% | theguardian.com 1.4% |
Between 2004 and 2026 the number of distinct domains cited grew from 363 to 3,685 — a 10× widening — while the top-10 share moved from 16% to 15.8%. By the Herfindahl–Hirschman index these pages are measurably less concentrated now than in 2004.
Every citation is classified against a curated list of platforms, publishers, vendors, academic and government domains. Roughly 78.2% of citations fall outside that list, so the shares below are quoted against all citations rather than against the classified subset — the honest denominator, and the reason this is two series rather than a pie.
Platform citations peak at 12.3% in 2005 and fall to 7%. Publisher citations rise from 4.5% to 12.2%. They cross in 2014, and publishers never fall behind again.
google.com was the single most-cited domain in 2004, at 4.9% of all citations. By 2026 it has dropped out of the top four entirely, displaced by nytimes.com. The press replaced the platform.
The encyclopedia did not stop covering search. It stopped leaning on the industry’s account of itself — including Google’s — and moved to journalism about it. Acceptance, in a reference work, means being written about rather than being the one writing.
Wikipedia is used here as an instrument, not a subject. An encyclopedia has to decide what words mean, and every one of those decisions is recorded, timestamped and signed: article genesis, redirect genesis, deletion, and the deletion debate itself — which survives in full even after the article it removed is gone. Because a reference work adopts a term only after it is established elsewhere, these dates measure acceptance, not invention.
The naming record runs from March 2001 to August 2026 — twenty-five years, because creation dates and redirect genesis exist from the beginning. The text corpus is 23 yearly snapshots covering 2004–2026, which is as far back as retrievable full-article captures go. Vocabulary measurements use the shorter span; genesis and naming findings use the longer one.
A term’s retention is its 2026 density as a share of its peak density — measured at the corpus’s last year, not at the term’s own last appearance. A phrase absent from the 2026 text retains 0%, with its last sighting reported separately. The distinction matters: measured at its own last nonzero year, an extinct phrase reads as fully alive — invisible text, unseen since 2005, would score 100%. Every retention figure on this page is recomputed from the raw yearly densities by an automated check before publication.
Wikipedia’s public deletion log begins 23 December 2004, so anything created and removed before that date is invisible to this method — which is why the 2003 article can only be called the earliest surviving attempt.
Vocabulary is measured on the text corpus, which starts in 2004; the 14 months between the article’s creation and the first retrievable snapshot are dated in the naming record but not measurable as prose. Domain classification is a curated list covering about 22% of citations, so platform and publisher shares are quoted against all citations rather than against the classified subset.
What Wikipedia deleted — the trade bodies, the conferences, the practitioners — is a substantial record in its own right and is deliberately not here. It is a question about institutions rather than about phrasing, and it is being written up separately.
The corpus is 6,083 yearly article snapshots pulled from Wikipedia’s own APIs and archived as Parquet on object storage; every n-gram measurement on this page is computed from that archive with DuckDB and committed as a versioned artifact before it renders. Our Wayback Machine MCP did the parts no live API can: recovering the bodies of deleted revisions and verifying what the founding citations serve today. Every chart is server-rendered SVG generated from those artifacts — no client JavaScript — and a verification harness recomputes all of the published figures from the raw artifacts on every change. Built with Claude.
SEO Bandwagon, "History of SEO Terms and Jargon on Wikipedia" (2026-08-16). https://seobandwagon.com/research/how-spamdexing-became-seo
This study came out of the same tooling we point at client sites — archives, revision histories and search data, read as evidence rather than anecdote.
Browse the SEO glossaryor talk to us · technical SEO audit · rank tracking · GBP benchmarks research