Skip to content

Improve technical SEO

You will improve the application’s initial HTML, metadata, and crawlable connections. You will audit canonical, robots, sitemap, HTTPS, and HTTP-response signals against a real public origin when one exists, or record those checks as Blocked without publishing or inventing a URL.

Level 1 improved page meaning for people and search systems. Technical SEO asks whether a crawler can discover a stable URL, request a useful response, read valid metadata and content, follow links, interpret crawl and index controls, and distinguish the preferred URL from duplicates. Each claim needs evidence from the correct layer.

What you will practice

  • Separate local source readiness, rendered-page evidence, public HTTP evidence, and search-engine index evidence.
  • Keep useful product purpose and navigation in the initial HTML instead of requiring an interaction to reveal all meaning.
  • Write valid page metadata and detect conflicting robots or canonical signals.
  • Use standard anchor elements with href values for discoverable navigation.
  • Explain the different roles and limits of canonical links, redirects, robots.txt, robots meta directives, and XML sitemaps.
  • Verify 200, redirect, and 404 behavior only against an authorized HTTP deployment.
  • Record crawl and index readiness without claiming indexing, traffic, or ranking.
  • New: Initial HTML response, rendered HTML, crawl eligibility, index eligibility, canonical URL, duplicate URL, HTTP status, redirect, robots.txt, robots meta directive, XML sitemap, public origin, and live URL inspection.
  • Reused: The tested application, visitor need, page title, meta description, headings, descriptive links, semantic HTML, valid source, responsive and keyboard checks, evidence statuses, Git commits, and quality-board boundaries.

Starting point

Before you start

  • The final tested application commit and interactive-test-report.md from the preceding lesson.
  • The private GitHub repository, quality-plan.md, and project board.
  • The source-and-Console issue closed with current validation evidence.
  • A teacher-approved public HTTPS origin if one already exists; publication is not required for this lesson.
  • VS Code, a current browser with DevTools, Git, and internet access for official references and optional live checks.
Current state
The application has interaction evidence, but its initial source, crawlable links, metadata, crawl controls, canonical decision, sitemap decision, and live response boundary have not been recorded as one technical SEO audit.
First action
Create technical-seo-report.md in the application repository and record whether an authorized public HTTPS origin exists.
First checkpoint
The report identifies the tested commit, publication state, canonical decision owner, source and live scope, evidence statuses, and the first local source check.
Help trigger
Use the recovery note or ask for help if no one can confirm the public origin, a live URL contains private data, source and rendered HTML disagree, a robots or canonical rule conflicts, a missing route returns 200, or a change requires hosting access you do not have.

You have completed the lesson when:

  • technical-seo-report.md identifies the starting commit, test environment, publication state, public origin or explicit absence, and scope;
  • every audit row contains source or live evidence, a final status, and a next action when it is not Pass;
  • index.html has valid head metadata with one accurate title and one page-specific meta description;
  • the initial HTML contains the product name, a useful purpose statement, primary control labels, and a visible explanation that does not depend on JavaScript execution;
  • internal navigation uses real a elements with resolvable href values;
  • JavaScript continues to own interaction state without replacing the page’s complete static meaning;
  • no unintended noindex, nofollow, duplicate canonical element, invalid head child, or blocked required CSS or JavaScript resource remains;
  • the source, rendered DOM, keyboard, narrow, zoom, and interaction regression checks pass after local changes;
  • when no approved public origin exists, canonical, robots, sitemap, redirect, 200, 404, and index checks remain Blocked with the missing condition and owner recorded;
  • when an approved public origin exists, the self-canonical, internal links, optional sitemap URLs, and redirect target use the same HTTPS origin and path policy;
  • a public robots.txt, when used, is available at the origin root, does not expose secrets, and does not block resources needed to render the page;
  • a public XML sitemap, when used, contains only preferred public URLs and is referenced consistently;
  • live checks record the actual 200, redirect, and missing-route response behavior without treating localhost as public proof;
  • no report claim promises crawling, indexing, a result appearance, traffic, or ranking;
  • the final tested commit is recorded in the technical SEO report and interactive regression evidence; and
  • the report and verified source changes are committed and pushed to the private repository.
Claim Required evidence Local source cannot prove
Metadata is present and valid Saved HTML and parsed document head What a search result will display
Content appears after JavaScript Rendered DOM after a clean load That every crawler renders it
URL returns 200 Response from the real public URL Index inclusion
Missing route returns 404 Response from a real missing public URL Quality of every error page
Crawler may request a path Public robots rules and public resource access That a crawler will visit it
Preferred URL is signaled Canonical, redirect, internal-link, and sitemap consistency Which canonical a search system will select
URL is indexed Search-engine index or URL-inspection evidence Ranking for a query

Localhost, file: URLs, preview servers, and private repositories are valuable development evidence. They are not public crawl or index evidence.

Add this structure to technical-seo-report.md:

technical-seo-report.md — starting structure
# Technical SEO report: Product name
## Tested build and environment
- Starting commit:
- Browser and version:
- Operating system:
- Test date:
- Run method:
## Publication boundary
- Public HTTPS origin: None, or exact approved origin
- Publication owner:
- Canonical URL policy: Blocked until origin exists, or exact policy
- Search Console property: None, or approved property
- Private data review: Pass, Issue, or Blocked
## Local source and render audit
Add the required local matrix.
## Public origin audit
Add the required live matrix. Use Blocked when no approved origin exists.
## Repairs and regressions
Record each source change and the repeated interaction checks.
## Final evidence
- Final tested commit:
- Local source result:
- Public origin result:
- Index evidence: Not tested, or exact approved tool evidence
- Open limits:
- Next action:

Use these statuses: Not run, Pass, Issue, Blocked, and Accepted limitation. Blocked is correct when a real origin, permission, server setting, or property does not exist.

Open the saved index.html in VS Code and use View page source in the browser. Do not use only the Elements panel: Elements shows the current DOM after parsing and JavaScript changes.

Add this local matrix to the report:

ID Check Expected result Actual result Status and evidence
HEAD-01 Parse the document head Only valid head elements; one title; one charset; one viewport Not run Not run
META-01 Read title and description Accurate, page-specific, useful without repeated phrases Not run Not run
HTML-01 Read initial main content Product name, purpose, controls, and explanation exist before JS Not run Not run
LINK-01 Inspect internal navigation Real anchors have resolvable href values and descriptive text Not run Not run
ROBOT-01 Search source for robots directives No unintended noindex or nofollow; decision is documented Not run Not run
CANON-01 Search source for canonical elements Zero until origin is confirmed, otherwise one correct absolute URL Not run Not run
RENDER-01 Compare source and rendered DOM Required static meaning persists; interaction content renders correctly Not run Not run
REG-01 Run focused interaction regression Primary, keyboard, status, narrow, and zoom paths still pass Not run Not run

Record the starting source before changing it. A clean-looking browser page does not show whether JavaScript supplied all useful text or whether the original head contains conflicting metadata.

Checkpoint: Local and live claims are separated

What now works
The report identifies the publication boundary and starting commit, contains the eight local checks, and uses Blocked rather than invented evidence for unavailable public conditions.
Files changed
technical-seo-report.md, index.html
What remains
Improve valid metadata, initial HTML meaning, and crawlable links, then rerun interaction regressions.
Next action
Inspect the saved head and correct HEAD-01 before writing new metadata.
If it does not work
If the public URL or canonical owner is uncertain, leave live configuration unchanged and record the exact person or decision needed. Continue with local source checks.

Keep metadata inside head. Adapt this reference to the product’s actual name and purpose:

index.html — local metadata boundary
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Project task tracker | Studio Practice</title>
<meta
name="description"
content="Plan, review, and update a short project task list with keyboard-accessible controls and visible completion status."
>
<link rel="stylesheet" href="styles.css">
<script src="app.js" defer></script>
</head>

The title identifies the specific product and its practice context. The description states what the page provides. Neither promises ranking or repeats a phrase for search manipulation.

Validate the complete saved HTML after editing. Invalid content inside head can change how later metadata is parsed. Keep visible headings and paragraphs in body, not inside head.

Do not add a keywords meta element, an index, follow robots element, or structured data because a checklist mentions them. Add a technical signal only when its role and evidence are clear.

Interactive records can come from JavaScript state. The product’s identity, purpose, controls, and explanation should not require a click or successful script execution to exist.

Adapt the existing main structure so it includes these static regions around the working application:

index.html — static meaning around the application
<main>
<header>
<p>Project planning practice</p>
<h1>Project task tracker</h1>
<p>
Add project tasks, update completion state, and review the current
completion summary.
</p>
<nav aria-label="Page sections">
<a href="#task-workspace">Use the task tracker</a>
<a href="#how-it-works">How the tracker works</a>
</nav>
</header>
<section id="task-workspace" aria-labelledby="task-workspace-title">
<h2 id="task-workspace-title">Task workspace</h2>
<!-- Keep the tested form, task list, and status region here. -->
</section>
<section id="how-it-works" aria-labelledby="how-it-works-title">
<h2 id="how-it-works-title">How the tracker works</h2>
<p>
The application validates task text, stores task records in page
memory, and renders the list and completion summary from that state.
Reloading restores the documented initial records.
</p>
</section>
</main>

Do not paste the comment in place of the application. Move the already tested form, list, and status elements into the task-workspace section without changing their IDs, labels, or event connections.

The two links are ordinary a elements with href values. JavaScript is not needed to discover or follow them. They describe sections on the same page, not separate indexable routes.

Compare source, no-script, and rendered states

Section titled “Compare source, no-script, and rendered states”

Run three checks:

  1. Source: View page source and find the product name, purpose, two links, control labels, and explanation.
  2. JavaScript unavailable: Disable JavaScript temporarily and reload. Confirm the product purpose and explanation remain readable. The dynamic task records can be unavailable because the page states that boundary.
  3. Rendered: Re-enable JavaScript, reload from the reset procedure, and run the primary interaction. Confirm the tested state-render behavior remains correct.

Technical SEO does not require duplicating dynamic records in static HTML. It requires an honest, useful initial response and a rendered result that does not contradict it.

Inspect every navigation element:

  • use a with href for navigation to a URL or fragment;
  • use button for actions that change current application state;
  • keep link text descriptive without its surrounding sentence;
  • confirm every local href resolves; and
  • do not use span, div, or onclick alone as a navigation link.

Search the saved source for noindex, nofollow, canonical, robots, and JavaScript that creates or changes those values.

For the intended public page:

  • an unintended noindex conflicts with index eligibility;
  • nofollow can affect link following and should have a specific reason;
  • several canonical elements create conflicting signals; and
  • JavaScript must not replace one source canonical with a different URL.

Do not remove an intentional privacy or staging control without approval. If the application must remain private, record that index eligibility is intentionally outside its publication state.

Checkpoint: The initial response explains the product

What now works
The valid head has specific metadata, initial HTML contains product meaning and real links, no unintended index control remains, and source, no-script, rendered, keyboard, narrow, zoom, and interaction checks agree.
Files changed
index.html, styles.css when layout needs a focused adjustment, app.js only when a confirmed integration defect needs repair, technical-seo-report.md
What remains
Audit real-origin canonical, robots, sitemap, redirect, success, missing-route, and index evidence when authorized.
Next action
Open Public origin audit and classify every live-only row before adding a canonical or crawler file.
If it does not work
If moving the tested controls breaks JavaScript, compare the required IDs and selectors with the last passing commit. Restore the DOM contract before changing SEO metadata.

Add these rows to the report:

ID Live-only check Expected result when public Actual result Status and evidence
PUBLIC-01 Request the approved HTTPS URL Public page is available without login or private data Not run Not run or Blocked
HTTP-01 Inspect the preferred page response Real origin returns 200 with the intended HTML Not run Not run or Blocked
HTTP-02 Request a clearly missing path Real origin returns 404, not a normal page with 200 Not run Not run or Blocked
REDIR-01 Request approved non-preferred variants Each configured variant redirects to one HTTPS policy Not run Not run or Blocked
CANON-02 Inspect the public page canonical One absolute self-canonical matches the preferred final URL Not run Not run or Blocked
ROBOTS-01 Request /robots.txt Intentional rules are public at the origin root and do not expose secrets Not run Not run or Blocked
MAP-01 Request the sitemap Valid XML lists only preferred public URLs Not run Not run or Blocked
INDEX-01 Inspect an approved search property URL state and selected canonical are recorded without a ranking claim Not run Not run or Blocked

When no approved public origin exists, set every row to Blocked with:

  • missing condition: approved public HTTPS origin;
  • owner: teacher, host administrator, or named product owner;
  • consequence: HTTP, canonical, crawler-file, and index evidence cannot be produced locally; and
  • next action: review these rows before the first authorized deployment.

Do not use https://example.com, localhost, a temporary preview URL, or a private GitHub file URL as the product’s canonical origin.

Add a canonical only after the origin is fixed

Section titled “Add a canonical only after the origin is fixed”

A canonical URL signals the preferred representative of duplicate or very similar URLs. It does not redirect a visitor, prevent crawling, guarantee selection, or create a public deployment.

When the publication owner confirms the exact final HTTPS URL, add one self-referential element to the saved HTML source:

Syntax example — replace before use
<link rel="canonical" href="https://www.example.com/task-tracker/">

www.example.com is a reserved syntax example. Do not leave it in product source. Replace the complete origin and path with the approved public URL, or omit the canonical.

Then confirm that:

  • the browser’s final address after redirects is the same URL;
  • internal links use the preferred scheme, host, case, path, and trailing-slash policy;
  • the sitemap uses that same URL when a sitemap exists;
  • no HTTP header or JavaScript supplies a different canonical; and
  • a duplicate variant redirects or otherwise uses a consistent canonical decision.

One wrong canonical can point search signals away from the product. An absent canonical is safer than a fabricated one.

Treat robots.txt as crawl guidance, not security

Section titled “Treat robots.txt as crawl guidance, not security”

robots.txt is a public file at the origin root. It manages crawler requests; it does not protect private content and is not the correct tool for canonical selection.

When the approved deployment needs a simple crawler file, adapt this public example:

robots.txt — public-origin example
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml

Replace the sitemap URL with the approved public origin. Omit the Sitemap line when no sitemap exists. Do not list a private route in order to hide it; the file itself can reveal the path.

Request /robots.txt from the real origin and record its response and complete rules. Confirm that required CSS, JavaScript, images, and public page paths are not blocked.

Create a sitemap only for real preferred URLs

Section titled “Create a sitemap only for real preferred URLs”

An XML sitemap can help a search system discover preferred URLs. It is a hint, not a guarantee of crawling or indexing.

For a real single-page deployment, the minimal file can be:

sitemap.xml — public-origin example
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/task-tracker/</loc>
</url>
</urlset>

Replace the example URL before publishing. Include only public canonical URLs that return success. Do not add fragment links such as #how-it-works, private routes, redirects, missing paths, search parameters, or invented future pages.

If the one-page host already generates an accurate sitemap, inspect that output instead of maintaining a second file. Record the source of truth and how it updates.

Verify public HTTP behavior when authorized

Section titled “Verify public HTTP behavior when authorized”

Complete this section only for an approved public origin. Use the browser Network panel or curl.exe on Windows to inspect the real response. Do not run live checks against an unrelated host.

  1. Request the preferred URL.

    Replace the placeholder with the approved address:

    Terminal window
    curl.exe -I https://REAL-APP-ORIGIN.example/approved-path/

    Record the complete URL, status, important headers, and test time. The preferred page must return 200 with the intended document.

  2. Request each approved non-preferred variant.

    Test only variants that the publication owner expects the host to support, such as an HTTP address or a different host name. Record every status and Location target. Each configured variant must finish at the one approved HTTPS URL without a loop.

  3. Request a clearly missing path.

    Use a path that cannot conflict with a real route:

    Terminal window
    curl.exe -I https://REAL-APP-ORIGIN.example/definitely-missing-seo-test

    The final response must be a real missing-resource status such as 404. A branded error page that returns 200 gives the wrong HTTP signal.

  4. Request crawler files that the deployment uses.

    Inspect /robots.txt and the exact sitemap URL. Record the status and body, not only whether the browser displays a file. If the deployment does not use one of these files, record the intentional absence and reason.

Local development-server responses can confirm local routing. They cannot replace evidence from the approved public origin.

Record index evidence without making a ranking claim

Section titled “Record index evidence without making a ranking claim”

Use a search-engine URL inspection tool only when the teacher or publication owner gives access to the correct property. Record:

  • the inspected URL;
  • the inspection date;
  • whether the URL is known or indexed;
  • the declared and selected canonical when the tool provides them; and
  • any crawl or index issue reported by the tool.

Do not request indexing, change a property, or submit a sitemap unless the owner authorizes that external action. If no approved property or access exists, set INDEX-01 to Blocked. An eligible page is not guaranteed to be crawled, indexed, displayed, or ranked.

Checkpoint: Public evidence is real or explicitly blocked

What now works
Each live-only row contains evidence from the approved origin or a Blocked status with the missing condition, owner, consequence, and next action. No example or local URL appears as product evidence.
Files changed
technical-seo-report.md, index.html, robots.txt and sitemap.xml only when the approved deployment uses them
What remains
Validate the source, rerun interaction regressions, resolve issues, and record the final commit.
Next action
Run the HTML validation and focused regression sequence from the tested build.
If it does not work
If a live result conflicts with the expected policy, preserve the response evidence and stop changing host settings. Assign the issue to the publication owner.

Technical SEO changes must not weaken the product. Validate the complete saved HTML, then repeat this focused regression sequence from the reset state:

  1. Submit one valid task and confirm the list, summary, status message, and focus result.
  2. Submit invalid or empty text and confirm the specific error and unchanged state.
  3. Complete or reopen one task and confirm the record and summary agree.
  4. Complete the primary path with the keyboard and visible focus.
  5. Check 320 CSS pixels and 200% zoom for reflow and reachable controls.
  6. Reload the documented starting state and confirm the Console has no new error.

Record the test IDs, environment, expected result, actual result, and evidence in the report. If a result changes, classify the defect, repair the smallest confirmed cause, and rerun both the failed check and the connected primary path.

After each repair, compare these layers again:

  • saved source;
  • JavaScript-unavailable page;
  • rendered DOM after a clean load;
  • real public response when available; and
  • index inspection only when authorized.

One passing layer does not make another layer pass.

For every row, replace Not run with Pass, Issue, Blocked, or Accepted limitation. Link or quote enough evidence to reproduce the result. Do not use a screenshot as the only evidence for an HTTP status, canonical value, source element, or response body.

Add a short final explanation under these headings:

State which links, page purpose, headings, metadata, and public URLs are available from the initial response. Separate local facts from public-origin facts.

Explain these roles in your own words:

  • a redirect changes the requested URL and sends the client to another URL;
  • a canonical identifies a preferred representative but does not redirect;
  • robots.txt controls crawler requests but does not protect private content;
  • a robots meta directive can restrict indexing of a page that the crawler can access;
  • a sitemap lists preferred public URLs as discovery hints; and
  • an HTTP status describes the response outcome.

Name the limits that apply to this product. For example, local source cannot prove a public response, a public response cannot prove indexing, and indexing cannot prove traffic or ranking.

The page looks complete only after JavaScript runs

Section titled “The page looks complete only after JavaScript runs”

Keep the product name, purpose, navigation, control labels, and explanatory text in the initial HTML. Keep live records and interaction state in JavaScript.

A canonical contains an example or preview URL

Section titled “A canonical contains an example or preview URL”

Remove it. Add a canonical only after the publication owner confirms the final public URL and path policy.

Do not use crawler guidance as access control. Remove private material from the public deployment and use real authorization where private access is required.

A missing route displays the home page with status 200

Section titled “A missing route displays the home page with status 200”

Preserve the response evidence and assign the server-routing issue to the publication owner. Do not mark HTTP-02 as passing because the page looks designed.

Mark the public-origin checks as Blocked. Local checks can still pass, and the lesson does not require publication.

Restore the last passing DOM contract, then reapply only the source change that has a clear purpose. Rerun the failed test and primary interaction path.

Self-check

Complete these checks against the required result.

  1. The report names the starting commit, environment, publication state, exact public origin or its absence, and decision owner.
  2. All eight local audit rows contain final evidence and a supported status.
  3. The saved head is valid and has one accurate title and one page-specific description.
  4. The initial HTML contains useful product meaning, control labels, explanation, and real anchor links.
  5. Source, no-script, rendered, keyboard, narrow, zoom, and interaction regression checks pass.
  6. No unintended robots directive, duplicate canonical, fabricated public URL, or blocked required rendering resource remains.
  7. All eight live-only rows use real-origin evidence or an explicit Blocked record with owner and next action.
  8. Any canonical, redirect, internal-link, robots, and sitemap signals use one approved URL policy.
  9. The report distinguishes crawl eligibility, index eligibility, index evidence, traffic, and ranking.
  10. The final report records the tested commit and does not claim an external result that the evidence cannot prove.

Complete the required result before choosing an extension. An extension does not change the definition of done.

Review the full diff before staging. Include robots.txt or sitemap.xml only when the approved public deployment uses the real files.

Terminal window
git status
git diff
git add index.html technical-seo-report.md
git diff --staged
git commit -m "Improve technical SEO readiness"
git push

Record the final commit hash in the report after the commit exists. If that requires one small follow-up commit, state which hash contains the tested source and which commit updates the record.

The next lesson defines an analytics question, configures a privacy-aware measurement boundary, verifies intentional events, and explains what the collected data can and cannot show.

If you stop, leave the repository in a committed state. Add the current audit row, evidence gathered, unresolved owner, next command or page to open, and the last passing regression result to technical-seo-report.md or the private project item.