Crawlers and SEO: how a page gets into search results, and HTML that ranks well
In the previous post, “CSR, SSR and SSG are about when the HTML gets assembled”, I sorted out CSR, SSR and SSG by where and when the HTML gets assembled. The crawler made a brief appearance there, but I never explained what a crawler actually is, or why being read by one is what gets a page into search results.
This post fills in that gap, in this order:
- What a crawler is
- What SEO is
- Why a page shows up in search results
- What kind of HTML tends to rank well (SEO in practice)
What a crawler is
A crawler is a program that goes around the web automatically and collects what is on each page. It moves from page to page by following links — crawling across the web — hence the name. Google’s crawler is called Googlebot, and Microsoft Bing’s is called Bingbot.
What it does is simple, repeated over and over:
- Request a URL and receive the HTML
- Read the page’s content, and its links to other pages, out of that HTML
- Add the links it found to its list of URLs to visit next
How a crawler finds pages
There are two main ways in.
- Links from other pages. If a page the crawler already knows links to yours, it follows that link and arrives at the new page.
- A sitemap. An XML file in which the site owner hands over a list saying “these are the URLs on this site”.
The usual way to tell crawlers where the sitemap lives is robots.txt, placed at the root of the site. robots.txt is a signpost for crawlers: it says what they may look at and where the sitemap is. This is the robots.txt for this blog:
User-agent: *
Allow: /
Sitemap: https://yomaru-tech.com/sitemap-index.xml
It means: “every crawler (*) may crawl every page (/), and the sitemap is here.”
Crawlers other than search engines
Search engines are not the only ones running crawlers. When you paste a URL into X or Slack, a preview with a title and image appears. That, too, is the result of the service’s crawler fetching the page and reading the tags in its HTML. This is why the previous post introduced the crawler as a program that “search engines and social networks send around the web automatically”.
Does a crawler run JavaScript?
In the previous post I wrote that the crawler “will not wait around for the cooking to finish”. That was a simplification for the sake of the analogy, so let me be more precise here.
Googlebot does have a step where it runs JavaScript and assembles the page (rendering). But this does not happen as soon as it receives the HTML. The page first goes into a queue and is rendered later. That wait can be a few seconds, or it can take longer.
On the other hand, many crawlers — such as the ones behind social media previews — do not run JavaScript at all and only read the HTML they first receive. When one of these reads a CSR page, it sees nothing but the empty plate from the previous post’s diagram.
In short, if the first HTML you return already has the content in it, every crawler can read it reliably, with no waiting. That is why SSR and SSG are said to be “good for search”.
What SEO is
SEO stands for Search Engine Optimization. It means shaping your pages and site so that people searching can find them more easily.
SEO has two broad sides.
| Side | What you do | Examples |
|---|---|---|
| Technical | Make sure crawlers read the page correctly | Put the content in the HTML, make links followable, publish a sitemap |
| Content | Write something useful to the person searching | Answer the question head-on, include concrete examples, keep information up to date |
SEO can sound like “the art of tricking search engines”, but it is really the opposite. Google itself publishes “make pages primarily for users, not for search engines” as its basic guideline, and tricks aimed at fooling search engines actually count against you (more on this later).
Why a page shows up in search results
Google describes how search works in three stages: crawling → indexing → serving search results.
Carrying on with the restaurant analogy from the previous post, it goes like this:
- Crawling: the crawler goes round the restaurants and photographs the food
- Indexing: the photos and notes it brought back are sorted and put into a restaurant guidebook
- Serving results: when a customer asks for “good ramen in Shibuya”, the right restaurants are picked out of the guidebook and shown in order of recommendation
The key point is that the search engine is not scouring the whole web at the moment you search. Results are picked from a guidebook built in advance — the index. In other words, a page shows up in search results because it is in the index.
Turn that around, and these pages will not show up:
- Pages nothing links to and that are not in a sitemap, so the crawler never finds them
- Pages the crawler finds but cannot read (the CSR plate from the previous post)
- Pages marked
noindex(covered below), which tells search engines not to index them
You can check whether a page is in Google’s index with the “URL Inspection” tool in Google Search Console, which Google provides for free.
What kind of HTML tends to rank well
To be clear up front: tidying up your HTML will not, on its own, push you up the rankings. The biggest factor is the content, and HTML is the means of getting that content across to crawlers without anything lost. With that said, here are the things that clearly make a difference.
1. Put the content in the first HTML
This is the previous post in a nutshell. With CSR, the HTML the server returns first is almost empty.
<!-- CSR: the first HTML returned. The content goes into #app only after the browser runs main.js -->
<body>
<div id="app"></div>
<script src="/main.js"></script>
</body>
The problem for crawlers is that the HTML returned here is the state before main.js has run. A crawler that does not run JavaScript can read nothing but an empty <div id="app"></div>.
With SSR or SSG, the content is there from the start.
<!-- SSR / SSG: the first HTML returned already has the content -->
<body>
<article>
<h1>CSR, SSR and SSG are about when the HTML gets assembled</h1>
<p>There are three answers to when and where a page's HTML gets assembled.</p>
</article>
</body>
Looking at this HTML, you cannot tell whether it came from SSR or SSG. The only difference is timing — whether the server assembles it on every request (SSR) or ahead of publishing (SSG) — and the HTML that comes back has the same shape either way. To a crawler, both are simply “HTML with the content already in it”.
For pages you want in search results, SSR or SSG is the safe choice.
2. Write a <title> and a <meta name="description">
<head>
<title>CSR, SSR and SSG are about when the HTML gets assembled</title>
<meta
name="description"
content="One axis — browser, server, or build time — is enough to sort out CSR, SSR and SSG."
/>
</head>
<title>becomes the large headline in search results. Make it different on every page and specific about what the page contains. If every page is titled “Home”, neither the searcher nor Google can tell what is inside.<meta name="description">may be used for the snippet shown under the title in search results. It does not directly raise your ranking, but it is what the searcher uses to decide whether to open the page.
As an aside, <meta name="keywords">, once widely used, is not used by Google for ranking. Writing it has no effect.
3. Use headings to show the document’s structure
<h1>Crawlers and SEO</h1>
<h2>What a crawler is</h2>
<h3>How a crawler finds pages</h3>
<h2>What SEO is</h2>
Use <h1> for the page’s subject, <h2> for the sections under it, <h3> for the parts under those, following the hierarchy. Crawlers read the headings to understand how the page is organised and what each part is about. Do not use headings for looks, as in “I want bigger text, so <h2>”. Adjust the look with CSS.
4. Write links as <a href>
<!-- ✗ the crawler cannot follow this -->
<span onclick="location.href='/en/posts/csr-ssr-ssg/'">previous post</span>
<!-- ✓ the crawler can follow this -->
<a href="/en/posts/csr-ssr-ssg/">previous post on the difference between CSR, SSR and SSG</a>
What Googlebot follows are links made with an <a> element that has an href attribute. An element that only switches the screen with JavaScript when clicked may behave like a link in the browser, but a crawler does not see it as one. The page it points to then goes undiscovered.
The link text (anchor text) should also describe what is on the other end, rather than “click here”. Crawlers use that text as a clue to what the linked page is about.
5. Give images an alt
<img src="/images/search-flow.png" alt="Diagram of the three stages: crawling, indexing and serving results" />
A crawler cannot see an image the way a person does, so it relies on the alt text to understand what the image shows. It is also used for image search, and it gets the content across to people using screen readers.
6. If the same page opens at several URLs, write a canonical
It is common for the same content to open at several URLs, such as https://example.com/posts/a/ and https://example.com/posts/a/?utm_source=x. Left alone, Google has to guess which one to show in search results.
<link rel="canonical" href="https://example.com/posts/a/" />
Use canonical to say “this is the representative URL for this page”, and the page’s standing is consolidated into that one URL. This blog outputs a canonical on every page.
7. For pages you do not want in search, write noindex
Some pages, like a “thanks for getting in touch” page, have no reason to be reached from search.
<meta name="robots" content="noindex" />
noindex is an instruction saying “do not put this page in the index”. If you also block crawling of that page in robots.txt, the crawler cannot read the page and never sees the noindex, so the bare URL can end up lingering in search results. When you want a page out of the index, keep crawling allowed and write noindex.
8. Use structured data to say what kind of page it is
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "BlogPosting",
"headline": "CSR, SSR and SSG are about when the HTML gets assembled",
"datePublished": "2026-09-06"
}
</script>
Structured data is information like “this is a blog post, this is its title, this is its publication date”, written in a fixed format (schema.org). Each type has its own set of fields: cooking time for a recipe, price and rating for a product, and so on.
When the conditions are met, the page may appear in search results with extras like rating stars or a price (a rich result). But adding it does not guarantee that display, and it does not directly raise your ranking.
9. Make it easy to read on a phone, and fast
Google mainly reads web pages with the smartphone version of Googlebot (mobile-first indexing), so pages that are hard to read on a phone are at a disadvantage. Start by putting this line in the <head> so the page fits the screen width:
<meta name="viewport" content="width=device-width, initial-scale=1" />
For speed, Google publishes three metrics called Core Web Vitals.
| Metric | What it measures |
|---|---|
| LCP (Largest Contentful Paint) | How long until the largest element on the page is displayed |
| INP (Interaction to Next Paint) | How long the screen takes to respond to an action such as a click |
| CLS (Cumulative Layout Shift) | How much the layout shifts around while loading |
SSG has the edge here too. Because it returns finished HTML as it is, the first view comes up quickly.
Things that backfire
Finally, what not to do. Google’s spam policies prohibit techniques such as:
- Keyword stuffing: repeating the same words unnaturally many times
- Hidden text: text people cannot see but crawlers can read, such as text the same colour as the background
- Cloaking: showing crawlers and people different content
These can get your rankings lowered, or get you removed from search results altogether.
Summary
- A crawler is a program that collects web pages by following links and sitemaps. Many crawlers do not run JavaScript, and even Googlebot runs it later, not straight away.
- SEO is the work of making your pages easier for searchers to find. It has a technical side (getting crawlers to read pages correctly) and a content side (being useful to people).
- A page shows up in search results because it is in the index. It goes through crawling → indexing → serving results.
- HTML that tends to rank well is HTML that gets the content across correctly. The basics carry most of the weight: put the content in the first HTML, get the
<title>and headings right, and write links as<a href>.
This is the reason behind the previous post saying “pages you want in search results should be SSR or SSG”. Search results are only ever picked from what crawlers have read and put in the index. So having the content in the first HTML is the foundation that every other SEO measure depends on.