Improve technical SEO
Outcome
Section titled “Outcome”You will improve the application’s initial HTML, metadata, and crawlable connections. You will audit canonical, robots, sitemap, HTTPS, and HTTP-response signals against a real public origin when one exists, or record those checks as Blocked without publishing or inventing a URL.
Why this matters
Section titled “Why this matters”Level 1 improved page meaning for people and search systems. Technical SEO asks whether a crawler can discover a stable URL, request a useful response, read valid metadata and content, follow links, interpret crawl and index controls, and distinguish the preferred URL from duplicates. Each claim needs evidence from the correct layer.
What you will practice
- Separate local source readiness, rendered-page evidence, public HTTP evidence, and search-engine index evidence.
- Keep useful product purpose and navigation in the initial HTML instead of requiring an interaction to reveal all meaning.
- Write valid page metadata and detect conflicting robots or canonical signals.
- Use standard anchor elements with href values for discoverable navigation.
- Explain the different roles and limits of canonical links, redirects, robots.txt, robots meta directives, and XML sitemaps.
- Verify 200, redirect, and 404 behavior only against an authorized HTTP deployment.
- Record crawl and index readiness without claiming indexing, traffic, or ranking.
What is new and what is reused
Section titled “What is new and what is reused”- New: Initial HTML response, rendered HTML, crawl eligibility, index eligibility, canonical URL, duplicate URL, HTTP status, redirect,
robots.txt, robots meta directive, XML sitemap, public origin, and live URL inspection. - Reused: The tested application, visitor need, page title, meta description, headings, descriptive links, semantic HTML, valid source, responsive and keyboard checks, evidence statuses, Git commits, and quality-board boundaries.
Starting point
Before you start
- The final tested application commit and interactive-test-report.md from the preceding lesson.
- The private GitHub repository, quality-plan.md, and project board.
- The source-and-Console issue closed with current validation evidence.
- A teacher-approved public HTTPS origin if one already exists; publication is not required for this lesson.
- VS Code, a current browser with DevTools, Git, and internet access for official references and optional live checks.
- Current state
- The application has interaction evidence, but its initial source, crawlable links, metadata, crawl controls, canonical decision, sitemap decision, and live response boundary have not been recorded as one technical SEO audit.
- First action
- Create technical-seo-report.md in the application repository and record whether an authorized public HTTPS origin exists.
- First checkpoint
- The report identifies the tested commit, publication state, canonical decision owner, source and live scope, evidence statuses, and the first local source check.
- Help trigger
- Use the recovery note or ask for help if no one can confirm the public origin, a live URL contains private data, source and rendered HTML disagree, a robots or canonical rule conflicts, a missing route returns 200, or a change requires hosting access you do not have.
Required result
Section titled “Required result”You have completed the lesson when:
technical-seo-report.mdidentifies the starting commit, test environment, publication state, public origin or explicit absence, and scope;- every audit row contains source or live evidence, a final status, and a next action when it is not
Pass; index.htmlhas validheadmetadata with one accuratetitleand one page-specific meta description;- the initial HTML contains the product name, a useful purpose statement, primary control labels, and a visible explanation that does not depend on JavaScript execution;
- internal navigation uses real
aelements with resolvablehrefvalues; - JavaScript continues to own interaction state without replacing the page’s complete static meaning;
- no unintended
noindex,nofollow, duplicate canonical element, invalidheadchild, or blocked required CSS or JavaScript resource remains; - the source, rendered DOM, keyboard, narrow, zoom, and interaction regression checks pass after local changes;
- when no approved public origin exists, canonical, robots, sitemap, redirect, 200, 404, and index checks remain
Blockedwith the missing condition and owner recorded; - when an approved public origin exists, the self-canonical, internal links, optional sitemap URLs, and redirect target use the same HTTPS origin and path policy;
- a public
robots.txt, when used, is available at the origin root, does not expose secrets, and does not block resources needed to render the page; - a public XML sitemap, when used, contains only preferred public URLs and is referenced consistently;
- live checks record the actual 200, redirect, and missing-route response behavior without treating localhost as public proof;
- no report claim promises crawling, indexing, a result appearance, traffic, or ranking;
- the final tested commit is recorded in the technical SEO report and interactive regression evidence; and
- the report and verified source changes are committed and pushed to the private repository.
Use the right evidence layer
Section titled “Use the right evidence layer”| Claim | Required evidence | Local source cannot prove |
|---|---|---|
| Metadata is present and valid | Saved HTML and parsed document head |
What a search result will display |
| Content appears after JavaScript | Rendered DOM after a clean load | That every crawler renders it |
| URL returns 200 | Response from the real public URL | Index inclusion |
| Missing route returns 404 | Response from a real missing public URL | Quality of every error page |
| Crawler may request a path | Public robots rules and public resource access | That a crawler will visit it |
| Preferred URL is signaled | Canonical, redirect, internal-link, and sitemap consistency | Which canonical a search system will select |
| URL is indexed | Search-engine index or URL-inspection evidence | Ranking for a query |
Localhost, file: URLs, preview servers, and private repositories are valuable development evidence. They are not public crawl or index evidence.
Create the technical SEO report
Section titled “Create the technical SEO report”Add this structure to technical-seo-report.md:
# Technical SEO report: Product name
## Tested build and environment
- Starting commit:- Browser and version:- Operating system:- Test date:- Run method:
## Publication boundary
- Public HTTPS origin: None, or exact approved origin- Publication owner:- Canonical URL policy: Blocked until origin exists, or exact policy- Search Console property: None, or approved property- Private data review: Pass, Issue, or Blocked
## Local source and render audit
Add the required local matrix.
## Public origin audit
Add the required live matrix. Use Blocked when no approved origin exists.
## Repairs and regressions
Record each source change and the repeated interaction checks.
## Final evidence
- Final tested commit:- Local source result:- Public origin result:- Index evidence: Not tested, or exact approved tool evidence- Open limits:- Next action:Use these statuses: Not run, Pass, Issue, Blocked, and Accepted limitation. Blocked is correct when a real origin, permission, server setting, or property does not exist.
Audit the initial source before editing
Section titled “Audit the initial source before editing”Open the saved index.html in VS Code and use View page source in the browser. Do not use only the Elements panel: Elements shows the current DOM after parsing and JavaScript changes.
Add this local matrix to the report:
| ID | Check | Expected result | Actual result | Status and evidence |
|---|---|---|---|---|
HEAD-01 |
Parse the document head | Only valid head elements; one title; one charset; one viewport | Not run | Not run |
META-01 |
Read title and description | Accurate, page-specific, useful without repeated phrases | Not run | Not run |
HTML-01 |
Read initial main content | Product name, purpose, controls, and explanation exist before JS | Not run | Not run |
LINK-01 |
Inspect internal navigation | Real anchors have resolvable href values and descriptive text | Not run | Not run |
ROBOT-01 |
Search source for robots directives | No unintended noindex or nofollow; decision is documented | Not run | Not run |
CANON-01 |
Search source for canonical elements | Zero until origin is confirmed, otherwise one correct absolute URL | Not run | Not run |
RENDER-01 |
Compare source and rendered DOM | Required static meaning persists; interaction content renders correctly | Not run | Not run |
REG-01 |
Run focused interaction regression | Primary, keyboard, status, narrow, and zoom paths still pass | Not run | Not run |
Record the starting source before changing it. A clean-looking browser page does not show whether JavaScript supplied all useful text or whether the original head contains conflicting metadata.
Checkpoint: Local and live claims are separated
- What now works
- The report identifies the publication boundary and starting commit, contains the eight local checks, and uses Blocked rather than invented evidence for unavailable public conditions.
- Files changed
technical-seo-report.md, index.html- What remains
- Improve valid metadata, initial HTML meaning, and crawlable links, then rerun interaction regressions.
- Next action
- Inspect the saved head and correct HEAD-01 before writing new metadata.
- If it does not work
- If the public URL or canonical owner is uncertain, leave live configuration unchanged and record the exact person or decision needed. Continue with local source checks.
Make the document head valid and specific
Section titled “Make the document head valid and specific”Keep metadata inside head. Adapt this reference to the product’s actual name and purpose:
<head> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1"> <title>Project task tracker | Studio Practice</title> <meta name="description" content="Plan, review, and update a short project task list with keyboard-accessible controls and visible completion status." > <link rel="stylesheet" href="styles.css"> <script src="app.js" defer></script></head>The title identifies the specific product and its practice context. The description states what the page provides. Neither promises ranking or repeats a phrase for search manipulation.
Validate the complete saved HTML after editing. Invalid content inside head can change how later metadata is parsed. Keep visible headings and paragraphs in body, not inside head.
Do not add a keywords meta element, an index, follow robots element, or structured data because a checklist mentions them. Add a technical signal only when its role and evidence are clear.
Keep useful meaning in initial HTML
Section titled “Keep useful meaning in initial HTML”Interactive records can come from JavaScript state. The product’s identity, purpose, controls, and explanation should not require a click or successful script execution to exist.
Adapt the existing main structure so it includes these static regions around the working application:
<main> <header> <p>Project planning practice</p> <h1>Project task tracker</h1> <p> Add project tasks, update completion state, and review the current completion summary. </p> <nav aria-label="Page sections"> <a href="#task-workspace">Use the task tracker</a> <a href="#how-it-works">How the tracker works</a> </nav> </header>
<section id="task-workspace" aria-labelledby="task-workspace-title"> <h2 id="task-workspace-title">Task workspace</h2> <!-- Keep the tested form, task list, and status region here. --> </section>
<section id="how-it-works" aria-labelledby="how-it-works-title"> <h2 id="how-it-works-title">How the tracker works</h2> <p> The application validates task text, stores task records in page memory, and renders the list and completion summary from that state. Reloading restores the documented initial records. </p> </section></main>Do not paste the comment in place of the application. Move the already tested form, list, and status elements into the task-workspace section without changing their IDs, labels, or event connections.
The two links are ordinary a elements with href values. JavaScript is not needed to discover or follow them. They describe sections on the same page, not separate indexable routes.
Compare source, no-script, and rendered states
Section titled “Compare source, no-script, and rendered states”Run three checks:
- Source: View page source and find the product name, purpose, two links, control labels, and explanation.
- JavaScript unavailable: Disable JavaScript temporarily and reload. Confirm the product purpose and explanation remain readable. The dynamic task records can be unavailable because the page states that boundary.
- Rendered: Re-enable JavaScript, reload from the reset procedure, and run the primary interaction. Confirm the tested state-render behavior remains correct.
Technical SEO does not require duplicating dynamic records in static HTML. It requires an honest, useful initial response and a rendered result that does not contradict it.
Audit crawlable links and index controls
Section titled “Audit crawlable links and index controls”Inspect every navigation element:
- use
awithhreffor navigation to a URL or fragment; - use
buttonfor actions that change current application state; - keep link text descriptive without its surrounding sentence;
- confirm every local
hrefresolves; and - do not use
span,div, oronclickalone as a navigation link.
Search the saved source for noindex, nofollow, canonical, robots, and JavaScript that creates or changes those values.
For the intended public page:
- an unintended
noindexconflicts with index eligibility; nofollowcan affect link following and should have a specific reason;- several canonical elements create conflicting signals; and
- JavaScript must not replace one source canonical with a different URL.
Do not remove an intentional privacy or staging control without approval. If the application must remain private, record that index eligibility is intentionally outside its publication state.
Checkpoint: The initial response explains the product
- What now works
- The valid head has specific metadata, initial HTML contains product meaning and real links, no unintended index control remains, and source, no-script, rendered, keyboard, narrow, zoom, and interaction checks agree.
- Files changed
index.html, styles.css when layout needs a focused adjustment, app.js only when a confirmed integration defect needs repair, technical-seo-report.md- What remains
- Audit real-origin canonical, robots, sitemap, redirect, success, missing-route, and index evidence when authorized.
- Next action
- Open Public origin audit and classify every live-only row before adding a canonical or crawler file.
- If it does not work
- If moving the tested controls breaks JavaScript, compare the required IDs and selectors with the last passing commit. Restore the DOM contract before changing SEO metadata.
Add the public-origin audit
Section titled “Add the public-origin audit”Add these rows to the report:
| ID | Live-only check | Expected result when public | Actual result | Status and evidence |
|---|---|---|---|---|
PUBLIC-01 |
Request the approved HTTPS URL | Public page is available without login or private data | Not run | Not run or Blocked |
HTTP-01 |
Inspect the preferred page response | Real origin returns 200 with the intended HTML |
Not run | Not run or Blocked |
HTTP-02 |
Request a clearly missing path | Real origin returns 404, not a normal page with 200 |
Not run | Not run or Blocked |
REDIR-01 |
Request approved non-preferred variants | Each configured variant redirects to one HTTPS policy | Not run | Not run or Blocked |
CANON-02 |
Inspect the public page canonical | One absolute self-canonical matches the preferred final URL | Not run | Not run or Blocked |
ROBOTS-01 |
Request /robots.txt |
Intentional rules are public at the origin root and do not expose secrets | Not run | Not run or Blocked |
MAP-01 |
Request the sitemap | Valid XML lists only preferred public URLs | Not run | Not run or Blocked |
INDEX-01 |
Inspect an approved search property | URL state and selected canonical are recorded without a ranking claim | Not run | Not run or Blocked |
When no approved public origin exists, set every row to Blocked with:
- missing condition: approved public HTTPS origin;
- owner: teacher, host administrator, or named product owner;
- consequence: HTTP, canonical, crawler-file, and index evidence cannot be produced locally; and
- next action: review these rows before the first authorized deployment.
Do not use https://example.com, localhost, a temporary preview URL, or a private GitHub file URL as the product’s canonical origin.
Add a canonical only after the origin is fixed
Section titled “Add a canonical only after the origin is fixed”A canonical URL signals the preferred representative of duplicate or very similar URLs. It does not redirect a visitor, prevent crawling, guarantee selection, or create a public deployment.
When the publication owner confirms the exact final HTTPS URL, add one self-referential element to the saved HTML source:
<link rel="canonical" href="https://www.example.com/task-tracker/">www.example.com is a reserved syntax example. Do not leave it in product source. Replace the complete origin and path with the approved public URL, or omit the canonical.
Then confirm that:
- the browser’s final address after redirects is the same URL;
- internal links use the preferred scheme, host, case, path, and trailing-slash policy;
- the sitemap uses that same URL when a sitemap exists;
- no HTTP header or JavaScript supplies a different canonical; and
- a duplicate variant redirects or otherwise uses a consistent canonical decision.
One wrong canonical can point search signals away from the product. An absent canonical is safer than a fabricated one.
Treat robots.txt as crawl guidance, not security
Section titled “Treat robots.txt as crawl guidance, not security”robots.txt is a public file at the origin root. It manages crawler requests; it does not protect private content and is not the correct tool for canonical selection.
When the approved deployment needs a simple crawler file, adapt this public example:
User-agent: *Allow: /
Sitemap: https://www.example.com/sitemap.xmlReplace the sitemap URL with the approved public origin. Omit the Sitemap line when no sitemap exists. Do not list a private route in order to hide it; the file itself can reveal the path.
Request /robots.txt from the real origin and record its response and complete rules. Confirm that required CSS, JavaScript, images, and public page paths are not blocked.
Create a sitemap only for real preferred URLs
Section titled “Create a sitemap only for real preferred URLs”An XML sitemap can help a search system discover preferred URLs. It is a hint, not a guarantee of crawling or indexing.
For a real single-page deployment, the minimal file can be:
<?xml version="1.0" encoding="UTF-8"?><urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <url> <loc>https://www.example.com/task-tracker/</loc> </url></urlset>Replace the example URL before publishing. Include only public canonical URLs that return success. Do not add fragment links such as #how-it-works, private routes, redirects, missing paths, search parameters, or invented future pages.
If the one-page host already generates an accurate sitemap, inspect that output instead of maintaining a second file. Record the source of truth and how it updates.
Verify public HTTP behavior when authorized
Section titled “Verify public HTTP behavior when authorized”Complete this section only for an approved public origin. Use the browser Network panel or curl.exe on Windows to inspect the real response. Do not run live checks against an unrelated host.
-
Request the preferred URL.
Replace the placeholder with the approved address:
Terminal window curl.exe -I https://REAL-APP-ORIGIN.example/approved-path/Record the complete URL, status, important headers, and test time. The preferred page must return
200with the intended document. -
Request each approved non-preferred variant.
Test only variants that the publication owner expects the host to support, such as an HTTP address or a different host name. Record every status and
Locationtarget. Each configured variant must finish at the one approved HTTPS URL without a loop. -
Request a clearly missing path.
Use a path that cannot conflict with a real route:
Terminal window curl.exe -I https://REAL-APP-ORIGIN.example/definitely-missing-seo-testThe final response must be a real missing-resource status such as
404. A branded error page that returns200gives the wrong HTTP signal. -
Request crawler files that the deployment uses.
Inspect
/robots.txtand the exact sitemap URL. Record the status and body, not only whether the browser displays a file. If the deployment does not use one of these files, record the intentional absence and reason.
Local development-server responses can confirm local routing. They cannot replace evidence from the approved public origin.
Record index evidence without making a ranking claim
Section titled “Record index evidence without making a ranking claim”Use a search-engine URL inspection tool only when the teacher or publication owner gives access to the correct property. Record:
- the inspected URL;
- the inspection date;
- whether the URL is known or indexed;
- the declared and selected canonical when the tool provides them; and
- any crawl or index issue reported by the tool.
Do not request indexing, change a property, or submit a sitemap unless the owner authorizes that external action. If no approved property or access exists, set INDEX-01 to Blocked. An eligible page is not guaranteed to be crawled, indexed, displayed, or ranked.
Checkpoint: Public evidence is real or explicitly blocked
- What now works
- Each live-only row contains evidence from the approved origin or a Blocked status with the missing condition, owner, consequence, and next action. No example or local URL appears as product evidence.
- Files changed
technical-seo-report.md, index.html, robots.txt and sitemap.xml only when the approved deployment uses them- What remains
- Validate the source, rerun interaction regressions, resolve issues, and record the final commit.
- Next action
- Run the HTML validation and focused regression sequence from the tested build.
- If it does not work
- If a live result conflicts with the expected policy, preserve the response evidence and stop changing host settings. Assign the issue to the publication owner.
Validate and rerun the interaction checks
Section titled “Validate and rerun the interaction checks”Technical SEO changes must not weaken the product. Validate the complete saved HTML, then repeat this focused regression sequence from the reset state:
- Submit one valid task and confirm the list, summary, status message, and focus result.
- Submit invalid or empty text and confirm the specific error and unchanged state.
- Complete or reopen one task and confirm the record and summary agree.
- Complete the primary path with the keyboard and visible focus.
- Check 320 CSS pixels and 200% zoom for reflow and reachable controls.
- Reload the documented starting state and confirm the Console has no new error.
Record the test IDs, environment, expected result, actual result, and evidence in the report. If a result changes, classify the defect, repair the smallest confirmed cause, and rerun both the failed check and the connected primary path.
After each repair, compare these layers again:
- saved source;
- JavaScript-unavailable page;
- rendered DOM after a clean load;
- real public response when available; and
- index inspection only when authorized.
One passing layer does not make another layer pass.
Complete the report
Section titled “Complete the report”For every row, replace Not run with Pass, Issue, Blocked, or Accepted limitation. Link or quote enough evidence to reproduce the result. Do not use a screenshot as the only evidence for an HTTP status, canonical value, source element, or response body.
Add a short final explanation under these headings:
What search systems can discover
Section titled “What search systems can discover”State which links, page purpose, headings, metadata, and public URLs are available from the initial response. Separate local facts from public-origin facts.
What the technical signals do
Section titled “What the technical signals do”Explain these roles in your own words:
- a redirect changes the requested URL and sends the client to another URL;
- a canonical identifies a preferred representative but does not redirect;
robots.txtcontrols crawler requests but does not protect private content;- a robots meta directive can restrict indexing of a page that the crawler can access;
- a sitemap lists preferred public URLs as discovery hints; and
- an HTTP status describes the response outcome.
What the evidence does not prove
Section titled “What the evidence does not prove”Name the limits that apply to this product. For example, local source cannot prove a public response, a public response cannot prove indexing, and indexing cannot prove traffic or ranking.
Common problems
Section titled “Common problems”The page looks complete only after JavaScript runs
Section titled “The page looks complete only after JavaScript runs”Keep the product name, purpose, navigation, control labels, and explanatory text in the initial HTML. Keep live records and interaction state in JavaScript.
A canonical contains an example or preview URL
Section titled “A canonical contains an example or preview URL”Remove it. Add a canonical only after the publication owner confirms the final public URL and path policy.
robots.txt blocks a private path
Section titled “robots.txt blocks a private path”Do not use crawler guidance as access control. Remove private material from the public deployment and use real authorization where private access is required.
A missing route displays the home page with status 200
Section titled “A missing route displays the home page with status 200”Preserve the response evidence and assign the server-routing issue to the publication owner. Do not mark HTTP-02 as passing because the page looks designed.
The live URL is unavailable
Section titled “The live URL is unavailable”Mark the public-origin checks as Blocked. Local checks can still pass, and the lesson does not require publication.
An SEO change breaks an interaction
Section titled “An SEO change breaks an interaction”Restore the last passing DOM contract, then reapply only the source change that has a clear purpose. Rerun the failed test and primary interaction path.
Self-check
Section titled “Self-check”Self-check
Complete these checks against the required result.
- The report names the starting commit, environment, publication state, exact public origin or its absence, and decision owner.
- All eight local audit rows contain final evidence and a supported status.
- The saved head is valid and has one accurate title and one page-specific description.
- The initial HTML contains useful product meaning, control labels, explanation, and real anchor links.
- Source, no-script, rendered, keyboard, narrow, zoom, and interaction regression checks pass.
- No unintended robots directive, duplicate canonical, fabricated public URL, or blocked required rendering resource remains.
- All eight live-only rows use real-origin evidence or an explicit Blocked record with owner and next action.
- Any canonical, redirect, internal-link, robots, and sitemap signals use one approved URL policy.
- The report distinguishes crawl eligibility, index eligibility, index evidence, traffic, and ranking.
- The final report records the tested commit and does not claim an external result that the evidence cannot prove.
Extension routes
Section titled “Extension routes”Complete the required result before choosing an extension. An extension does not change the definition of done.
Official references
Section titled “Official references”- Google Search technical requirements
- Google Search JavaScript SEO basics
- Google Search crawlable links
- Google Search canonical guidance
- Google Search robots.txt introduction
- Google Search sitemap guidance
Commit the completed audit
Section titled “Commit the completed audit”Review the full diff before staging. Include robots.txt or sitemap.xml only when the approved public deployment uses the real files.
git statusgit diffgit add index.html technical-seo-report.mdgit diff --stagedgit commit -m "Improve technical SEO readiness"git pushRecord the final commit hash in the report after the commit exists. If that requires one small follow-up commit, state which hash contains the tested source and which commit updates the record.
Next lesson
Section titled “Next lesson”The next lesson defines an analytics question, configures a privacy-aware measurement boundary, verifies intentional events, and explains what the collected data can and cannot show.
Safe resume point
Section titled “Safe resume point”If you stop, leave the repository in a committed state. Add the current audit row, evidence gathered, unresolved owner, next command or page to open, and the last passing regression result to technical-seo-report.md or the private project item.