MarTechQuick
Beginner

Technical SEO Fundamentals

14 min read

Learn
Quick Reading
Estimated 14 mins
Prereq
Foundational
No experience required
Interactive
Static Playbook
Static guide & reference tables

SEO starts with access, not keywords

Before a page can rank, a search engine has to be able to find it, fetch it, and be allowed to show it. That three-step chain — crawlability, indexation, and rendering — is technical SEO, and it fails silently more often than people expect. A page can have perfect copy and still rank nowhere because a single noindex tag, a misconfigured redirect, or a blocked resource in robots.txt quietly removed it from consideration.

Technical SEO is the engineering layer underneath content and links. It doesn't make a page more persuasive; it makes a page *eligible* to compete at all.

Crawling vs. indexing: two different gates

These terms get conflated constantly, but they're separate systems with separate failure modes:

- Crawling is Googlebot fetching a URL. Controlled by robots.txt (a request, not a lock — it stops fetching, not indexing) and by internal linking (unlinked pages are rarely discovered).
- Indexing is Google deciding to store the page in its database as a retrievable candidate. Controlled by the <meta name="robots" content="noindex"> tag or the X-Robots-Tag HTTP header, and gated by perceived quality — thin or duplicate pages get crawled but never indexed.

A page can be crawled but not indexed (common for low-value pages), and in rare cases indexed without being crawled (Google indexes a URL from a link's anchor text alone if it's blocked by robots.txt, showing it with no snippet). Confusing the two leads teams to 'fix' the wrong control.

robots.txt does not remove pages from the index

A common and costly mistake: blocking an already-indexed URL in robots.txt to try to remove it from search results. This does the opposite of what's intended — Googlebot can no longer crawl the page to see the noindex tag, so the URL can persist in the index indefinitely, now showing as a bare URL with no description. To deindex a page, use noindex (or a 410/404) and let it be crawled at least once, then optionally block it afterward.

robots.txt: the crawl gatekeeper

robots.txt sits at the domain root and issues crawl directives per user-agent. It's plain text, cached by search engines for roughly 24 hours, and it fails open — a missing file means 'crawl everything,' and a syntax error can accidentally block the entire site.

robots.txt
text
User-agent: *
Disallow: /admin/
Disallow: /checkout/
Disallow: /*?sort=          # block crawlable but non-canonical filtered URLs
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml

Canonical tags: telling search engines which URL is real

Most sites unintentionally generate multiple URLs for the same content — via tracking parameters (?utm_source=), session IDs, faceted navigation filters, or http vs https and www vs non-www variants. Without guidance, search engines have to guess which version to index, and they'll often split ranking signals across duplicates instead of consolidating them onto one strong page.

The rel="canonical" link tag, placed in the <head>, tells crawlers: 'this URL is a variant; treat this other URL as the authoritative one.' It's a strong hint, not a directive — Google can and does override a canonical it disagrees with, usually when the canonical target isn't actually equivalent content.

canonical_example.html
text
<!-- On https://example.com/shoes?color=red&utm_source=newsletter -->
<link rel="canonical" href="https://example.com/shoes" />

Core Web Vitals: performance as a ranking input

Google measures real-world page experience through three Core Web Vitals metrics, collected from actual Chrome users (field data via the Chrome UX Report), not synthetic lab tests alone:

- LCP (Largest Contentful Paint) — time until the largest visible element renders. Target: under 2.5s. Usually bottlenecked by slow server response, render-blocking CSS/JS, or unoptimized hero images.
- INP (Interaction to Next Paint) — responsiveness to user input, replacing First Input Delay in 2024. Target: under 200ms. Usually bottlenecked by long JavaScript tasks on the main thread.
- CLS (Cumulative Layout Shift) — visual stability; how much content jumps around as it loads. Target: under 0.1. Usually caused by images/ads without reserved dimensions or web fonts causing layout reflow (FOIT/FOUT).

Core Web Vitals thresholds

MetricGoodNeeds ImprovementPoor
LCP≤ 2.5s2.5s – 4.0s> 4.0s
INP≤ 200ms200ms – 500ms> 500ms
CLS≤ 0.10.1 – 0.25> 0.25

Site architecture: crawl budget and link equity

Large sites don't get infinite crawling — Google allocates a crawl budget per site based on server capacity and perceived value, and every wasted crawl on a low-value URL (infinite faceted-filter combinations, expired listing pages, internal search results) is a crawl not spent on pages you want ranked.

A flat, shallow architecture — ideally reaching any page within 3-4 clicks from the homepage — distributes internal link equity (PageRank) efficiently and helps crawlers discover new or updated pages quickly. Orphan pages (no internal links pointing to them) are the most common architecture failure: technically live, but functionally invisible unless they're in the sitemap and get crawled through it alone.

XML sitemaps are a hint list, not a guarantee

Submitting a URL in sitemap.xml increases the chance it gets crawled and considered, but it is not a command to index. Sitemaps matter most for large or poorly-linked sites where discovery through internal links alone is unreliable — they're a discovery aid, not a ranking lever.

What's next

Technical SEO makes pages eligible to rank; keyword research decides which pages are worth making eligible for in the first place, and how to prioritize the build queue.

Next: Keyword Research Framework →

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.