SEO glossary · Search fundamentals · Guide
Search engine basics: how search engines work
A software system that crawls, indexes and ranks web content, then returns matches for whatever a person types or speaks as a query.
That sentence is the glossary's definition of a search engine. This page is the long version of it: what happens between a crawler finding your page and a results page being assembled, in plain English, and what each stage asks of a small business's site.
Every term on the page links to its entry in the SEO glossary, and the definitions are the glossary's own, imported rather than retyped. Every date and figure carries a numbered source: the owner's own documentation where one exists, or research published on this site.
About this page
- Author
- Kyle Alm
- Definitions
- The SEO glossary's own, imported at build time
- Numbers
- From the glossary data and the research pages in the sources, never typed in
- Terms linked
- 61
- Sources
- 60
What is a search engine?
A search engine is two things joined together: an index of pages it has already fetched, and the systems that decide which of those pages to show for a query. MDN Web Docs describes the work in three steps: crawling the web by following links, indexing what the crawl finds, and searching that index to find and rank pages for a query[1].
It is not a browser. A browser such as Chrome, Safari, Firefox or Edge is the program on your device that fetches and displays pages; a search engine is a service you reach through it[2][1]. The two blur together because the browser's address bar doubles as a search box, sending anything that isn't a web address to whichever search engine is set as the default[3].
It is not a directory either. Yahoo said in 2014 that it had begun nearly twenty years earlier as a human-edited directory of sites, and it retired that directory on December 31, 2014[4]; DMOZ, the volunteer-edited Open Directory Project, closed on March 17, 2017[5]. A web directory lists what editors chose to file. A search engine finds pages on its own.
Google likens its index to “the index in the back of a book,” with an entry for every word on every page it holds, puts its size at “well over 100,000,000 gigabytes,” and says crawlers build most of it[6].
How do search engines work?
Google describes Search as three stages, crawling, indexing and serving, and says that not every page makes it through each one[7].
Crawling. A crawler such as Googlebot discovers a URL, mostly by following links from pages it already knows, then downloads the page's text, images and other files[7][8]. It doesn't fetch everything it discovers: robots.txt can disallow a URL, and a page behind a login can't be reached[7]. Googlebot reads only the first 2 MB of a page[8]. Two files steer this stage. Robots.txt is the plain-text file at a site's root, defined by RFC 9309, that tells crawlers which URLs they may request[9][10]. A sitemap lists a site's pages so a search engine can discover them without depending on links alone; Google says a sitemap doesn't guarantee that every listed URL gets crawled or indexed[11].
Indexing. Google analyzes what it fetched: the text, key tags such as the title element, and whether the page duplicates another[7]. Choosing one representative URL from a set of duplicates is canonicalization; Google crawls the canonical most often and treats a rel="canonical" preference as a hint rather than a rule[12]. A crawled page can still stay out of the index: a noindex rule keeps it out, and Google says indexing isn't guaranteed even for pages it has crawled[13][7]. What this stage builds is the reverse index the original Google paper described, which maps each word to the documents containing it, so a query can jump straight to candidate pages[14].
Serving. When someone searches, Google's ranking systems look the query up in the index and order the matches, which is the subject of the next section[7]. Google also personalizes some results, for example by location or search history, so two people can see the same page at different positions for the same words[15][16].
Bing runs the same three stages with its own crawler, Bingbot, and its own index[17].
What decides the order of results?
Google's explanation of ranking names five kinds of signal: the meaning of the query, the relevance of the content, the quality of the content, the usability of the page, and the searcher's context and settings[18].
- Meaning. Google builds language models to work out what a query intends, and a synonym system can match a page that doesn't contain the exact words used[18]. RankBrain, launched in 2015, and BERT, applied to Search from October 2019, are two of the named systems that do this work[15][19][20].
- Relevance. Google calls a page containing the same keywords as the query “the most basic signal that information is relevant,” and says it also uses aggregated, anonymized interaction data to estimate relevance[18].
- Quality. Links are one input. PageRank, which scores a page by treating each link to it as a vote weighted by the linking page's standing, was one of Google's core ranking systems at launch and is still part of them, as one signal among many[15]. Google's quality raters assess experience, expertise, authoritativeness and trust, but Google says E-E-A-T isn't itself a ranking factor[21].
- Usability. Google's ranking systems use Core Web Vitals, its three page experience metrics for loading, responsiveness and visual stability, and Google says it still seeks to show the most relevant content even when page experience is sub-par[22][23].
- Context. Location, language, device and past searches all shape what is shown, and personalization changes the order of some results rather than all of them[16][18].
Freshness is a signal of its own for queries where recent information is expected, such as breaking news; Google's guide lists freshness systems among its ranking systems[15].
Local results are a separate ranking. Google says local results are based mainly on relevance, distance and prominence, that it uses a searcher's location even when the query names no place, and that a business cannot pay for or request a better local position[24][25]. Google bases prominence partly on how many websites link to a business and how many reviews it has[24].
Our own measurement of that last point is narrower than the sales pitch[26]: Review count does not predict map-pack rank. The most-reviewed business in a market ranks 10th on average, and top-3 businesses out-review their nearest rivals in only 47.5% of head-to-head matchups — worse than a coin flip. Reviews behave as a visibility threshold (~75 reviews), not a ladder.
Spam is the other side of ranking. Google's spam policies define techniques that deceive users or manipulate its systems, its automated systems catch most violations, and a human reviewer's manual action can demote a site or remove it from results[27].
What is on a search results page?
A results page, the SERP, is assembled from blocks, and Google's Visual Elements gallery catalogs them[28]. The ones a small business meets most:
- Organic results. Free listings that appear because they are relevant to the query; Google Ads' glossary defines them that way and contrasts them with ads[29]. Google says it takes no money to include or rank a site in these results[30]. This is what SEO competes for.
- Ads. Paid placements marked Sponsored, priced and positioned by an auction that runs every time someone searches[31][29].
- AI Overviews and AI Mode. An AI-generated summary with links to supporting pages, shown only when Google's systems judge it adds to the classic results, and a separate mode for complex questions[32]. Both use query fan-out, several related searches run behind the scenes to gather source material[33]. A page needs no special file or markup to be linked from one, only to be indexed and eligible for a snippet[32].
- Featured snippets and People Also Ask. A featured snippet reverses the usual order and shows a page's descriptive text first; People Also Ask is a group of related questions, each expanding to a snippet-style answer[34][28].
- The local pack. A block of nearby businesses' Google Business Profiles for searches with local intent, ranked on relevance, distance and prominence[24].
- Knowledge panels. Boxes about a person, place or thing in Google's Knowledge Graph, which Google introduced on May 16, 2012, as a move from matching strings toward understanding things[35].
- Sitelinks, related searches and the rest. Google clusters extra links from the same site under a result when it judges them useful, generates related searches automatically, and has blended images, video, news and maps into one page since the universal search change it announced on May 16, 2007[36][28][37].
Rich results, listings with extra visual or interactive elements, come from structured data in a page's markup, and Google adds and removes types on its own schedule: it retired FAQ rich results on May 7, 2026[38][39].
The vocabulary is younger than the blocks. In our study of Wikipedia's record, the phrase “search engine results page” first appears in the corpus in 2006, peaks in 2024, and keeps 96% of that peak usage today[40].
What does a search engine need from your site?
Google publishes three technical requirements for a page to be eligible for its index: Googlebot isn't blocked, the page works, meaning the server answers with a success status, and the page has indexable content[41]. Google says most sites pass them without their owners noticing[42]. Past that line, the job is to be understood, relevant and trusted.
- Reachable. A working page at a stable URL, a robots.txt that doesn't block what you want found, and a sitemap for the pages that links alone might miss[41][10][11]. One trap: a page blocked by robots.txt can still show up in results if other sites link to it, because Google never crawls it to see a noindex rule[10].
- Understood. The title element is what the HTML Standard defines as a page's title, and Google said in 2021 that it used title text for the result headline more than 80% of the time[43][44]. The meta description can become the snippet and, per a 2007 Google post, an accurate one can improve clickthrough without affecting ranking[45]. Every page you care about should be linked from at least one other page on your site, with link text that describes the destination[46]. Structured data in schema.org vocabulary helps Google classify a page and can earn a rich result[38].
- Relevant. Match the intent, not just the words. Google's rater guidelines sort queries into Know, Do, Website and Visit-in-Person intents[47]. A page that meets a Know query with a sales pitch, or a Visit-in-Person query with an essay, is the wrong page type whatever its keywords.
- Trusted. Backlinks, links from pages on other sites, feed PageRank, which remains one signal among many[15][48]. Google's spam policies name link spam, keyword stuffing and cloaking among the ways to lose that trust[27].
- Local. A Google Business Profile is the free listing behind the local pack, which Google compiles from owner edits, user contributions, crawled web content and licensed data[49]. In our benchmark of 164,241 Maps-visible businesses across 54 metros, 15.6% had no website field at all, a share of businesses that already rank, not of all businesses[50].
The free tools on this site apply these basics to a real page: the free SEO audit and the Chrome extension check a page against them, and the rank tracker reads where a page stands in Google's results.
How did search engines get here?
Larry Page and Sergey Brin built BackRub at Stanford, ranking pages by the links between them; they renamed it Google, and Google Inc. was incorporated in 1998[51]. Their 1998 paper described the inverted index and PageRank, and Google says PageRank is still part of its core ranking systems[14][15].
Wikipedia's record, read in our study of 5,901,821 words across 365 articles from 2004 to 2026, shows the abuse was named before the profession: the encyclopedia's article on spamdexing dates from March 10, 2002, eleven and a half months before its article on search engine optimization was created on February 24, 2003[40][52][53].
Directories gave way to crawling: Yahoo retired its directory on December 31, 2014, and DMOZ closed on March 17, 2017[4][5].
Google announced universal search on May 16, 2007[37], and the Knowledge Graph on May 16, 2012[35]. Its language systems followed: Hummingbird in August 2013, RankBrain in 2015, BERT from October 2019, and MUM in May 2021[15][19][20][54].
Page experience entered the signals: Google began using HTTPS as a lightweight ranking signal in August 2014[55], and replaced First Input Delay with Interaction to Next Paint as a Core Web Vital in March 2024[56].
Generative AI arrived as Search Generative Experience, a Labs experiment announced in May 2023, became AI Overviews on May 14, 2024, and AI Mode reached Search Labs on March 5, 2025[57][39].
The words changed with the machinery. Of the 640 terms in the SEO glossary, 233 carry a date from 23 years of Wikipedia revisions, so you can see which of these words the industry still uses and which it has dropped.
FAQs
Do I need both a browser and a search engine?
How do search engines work step by step?
Crawl, index, serve: a crawler fetches pages by following links, the search engine analyzes and stores them in its index, and ranking systems pick and order matches for each query[7]. A page can drop out at any step: blocked by robots.txt, kept out by a noindex rule, or ranked below the first page[10][13].
What is the difference between Google Search and Google?
Is Google the only search engine?
No. Bing runs its own crawler, Bingbot, and its own index[17], and AI answer engines fetch pages with crawlers of their own: OpenAI's OAI-SearchBot for ChatGPT search and PerplexityBot for Perplexity[58][59]. The same basics apply to all of them: a page has to be reachable, readable and clear about what it is.
How do I learn SEO as a beginner?
Start with the three stages on this page, then Google's own SEO Starter Guide, which says there are no secrets that automatically rank a site first[60]. Look up each term as you meet it in the SEO glossary, and run the free SEO audit on your own site to see the basics applied to real pages.
Terms on this page
Every term this guide uses, linked to its glossary entry.
Sources
- MDN Web Docs, “Search engine.” (documentation)
- MDN Web Docs, “Browser.” (documentation)
- Google Chrome Help, “Set your default search engine & site search shortcuts.” (primary)
- Yahoo, via Internet Archive Wayback Machine, “Progress Report: Continued Product Focus.” Archived capture, September 26, 2014. (primary)
- DMOZ, via Internet Archive Wayback Machine, “DMOZ homepage.” Archived capture, March 19, 2017. (primary)
- Google (How Search Works), “How Google Search organizes information.” (primary)
- Google Search Central, “In-depth guide to how Google Search works.” (primary)
- Google Search Central, “Googlebot.” (primary)
- IETF (RFC Editor), “RFC 9309: Robots Exclusion Protocol.” (standard)
- Google Search Central, “Introduction to robots.txt.” (primary)
- Google Search Central, “Learn about sitemaps.” (primary)
- Google Search Central, “What is canonicalization.” (primary)
- Google Search Central, “Block Search indexing with noindex.” (primary)
- Sergey Brin and Lawrence Page (Stanford University), “The Anatomy of a Large-Scale Hypertextual Web Search Engine.” (research)
- Google Search Central, “A guide to Google Search ranking systems.” (primary)
- Google Search Help, “Personalization & Google Search results.” (primary)
- Bing Webmaster Tools, “Overview of Bing crawlers (user agents).” (primary)
- Google (How Search Works), “Automatically generating and ranking results.” (primary)
- Pandu Nayak, Google (The Keyword), “How AI powers great search results.” (primary)
- Pandu Nayak, Google (The Keyword), “Understanding searches better than ever before.” (primary)
- Google Search Central, “Creating helpful, reliable, people-first content.” (primary)
- Google Search Central, “Understanding Core Web Vitals and Google search results.” (primary)
- Google Search Central, “Understanding page experience in Google Search results.” (primary)
- Google Business Profile Help, “Tips to improve your local ranking on Google.” (primary)
- Google Search Help, “Understand & manage your location when you search on Google.” (primary)
- SEO Bandwagon, “The Periodic Table of SEO Ranking Factors.” Review count, measured; last verified August 28, 2026. (research)
- Google Search Central, “Spam policies for Google web search.” (primary)
- Google Search Central, “Visual Elements gallery of Google Search.” (primary)
- Google Ads Help, “Organic search result.” (primary)
- Google Search Central, “Do you need an SEO?” (primary)
- Google Ads Help, “Auction.” (primary)
- Google Search Central, “AI features and your website.” (primary)
- Google Search Central, “Optimizing your website for generative AI features on Google Search.” (primary)
- Google Search Central, “Featured snippets and your website.” (primary)
- Google (The Keyword), “Introducing the Knowledge Graph: things, not strings.” (primary)
- Google Search Central, “Sitelinks.” (primary)
- Official Google Blog, “Universal search: The best answer is still the best answer.” (primary)
- Google Search Central, “Introduction to structured data markup in Google Search.” (primary)
- Google Search Central, “Latest documentation updates.” (primary)
- SEO Bandwagon, “History of SEO Terms and Jargon on Wikipedia.” (research)
- Google Search Central, “Google Search technical requirements.” (primary)
- Google Search Central, “Google Search Essentials.” (primary)
- Google Search Central, “Influencing Title Links in Google Search.” (primary)
- Google Search Central Blog, “An update to how we generate web page titles.” (primary)
- Google Search Central Blog, “Improve snippets with a meta description makeover.” (primary)
- Google Search Central, “Link best practices for Google.” (primary)
- Google (Search Quality Rater Guidelines), “General Guidelines (Search Quality Rater Guidelines), 12.7 Understanding User Intent.” (primary)
- Search Console Help, “Links report.” (primary)
- Google Business Profile Help, “Understand how Google sources & uses info in Business Profiles & local search results.” (primary)
- SEO Bandwagon, “Google Business Profile Benchmarks 2026.” Edition: week of Sep 8, 2026. (research)
- About Google, “Our story: From the garage to the Googleplex.” (primary)
- Wikipedia, “Spamdexing, first revision.” March 10, 2002. (primary)
- Wikipedia, “Search engine optimization, first revision.” February 24, 2003. (primary)
- Pandu Nayak, Google (The Keyword), “MUM: A new AI milestone for understanding information.” (primary)
- Google Search Central Blog, “HTTPS as a ranking signal.” (primary)
- Google Search Central Blog, “Introducing INP to Core Web Vitals.” (primary)
- Google (The Keyword), “Supercharging Search with generative AI.” (primary)
- OpenAI, “Overview of OpenAI Crawlers.” (primary)
- Perplexity, “Perplexity Crawlers.” (primary)
- Google Search Central, “Search Engine Optimization (SEO) Starter Guide.” (primary)
Link to this guide
To cite this page on your own site, paste this HTML:
<a href="https://seobandwagon.com/search-engine-basics">Search engine basics: how search engines work</a>Ready to get found on Google?
Run the free SEO audit to see how your own site handles each stage.