You already know title tags, headings, and meta descriptions matter. Those get their own deep dives elsewhere. This is the list for everything else: the smaller pieces of markup that show up in every client audit, the ones that get argued about in briefs and Slack threads without anyone checking whether Google still cares. Some of these html tags for seo still do real work. A few never did, and one or two are pure legacy weight that somebody copied into a template years ago and nobody has removed since.
"Matters for SEO" does not mean "is a ranking factor" here, and that distinction is the whole point of this list. A robots directive can pull a page out of results entirely. A canonical tag can shift which URL Google shows for a set of duplicates. Structured data can make a page eligible for a search feature it would not otherwise get. None of that is the same as a keyword boost, and treating it that way is how audits end up wasting a client's dev hours on the wrong fix.
Here is the fast reference, then each item with what it actually does and the move to make on a real site. Think of it as the shortlist of seo html elements worth an agency's time once the bigger pieces are handled.
| Markup | Verdict | What it actually does |
|---|---|---|
| Robots meta directives | High priority when indexing or previews need control | Tells a crawler whether a page may be indexed and how it may be shown |
| Canonical link element | High priority wherever duplicate URLs exist | States a preferred URL among near-duplicates; Google can still choose differently |
| Image alt attribute | Worth doing on meaningful images | Describes an image for accessibility and helps Google understand it |
| Structured data markup | Conditional, valuable for eligible page types | Describes page content and can qualify it for a search feature, not guaranteed to appear |
| Hreflang link elements | Essential only for real locale alternatives | Points to the corresponding version of a page in another language or region |
| og:image metadata | Conditional display control, not a ranking lever | Can inform which thumbnail Google shows in Search and Discover |
| Valid head and semantic HTML | Foundational hygiene | Keeps metadata readable and content accessible to crawlers |
| Meta keywords and old scheduling tags | Drop from the checklist | Do nothing for Google indexing or ranking |
Robots meta directives
A page-level tag like <meta name="robots" content="noindex"> sits in the document head and tells Google not to show that page in results, once Google has actually fetched and processed the directive. That last part trips people up constantly. If the URL is blocked in robots.txt, Google may never see the on-page noindex at all, so the two controls are not interchangeable and blocking a page in robots.txt is not the same as removing it from the index.
The defaults are index and follow, so writing index,follow on every page does nothing useful. Where this tag earns its place is preview control: nosnippet, max-snippet, max-image-preview, and the text-level data-nosnippet attribute all shape what Google may show from a page without touching whether the page ranks. A googlebot-specific version of the tag addresses Google only, and when directives conflict, the more restrictive one wins.
The agency move here is simple and worth doing before anything else on this list: check the rendered head of a page the client believes is indexable, and confirm there is no accidental noindex or an overly restrictive snippet directive left over from a staging build. This is also the family Google points to when a publisher wants to limit what shows up in its AI-driven search features, so it is worth checking on pages a client wants kept out of AI Overviews. It does not extend to controlling every other AI engine, and you should not tell a client it does.
Canonical link element
The canonical link element, <link rel="canonical" href="[preferred URL]">, states which URL you would rather Google treat as the representative one when several pages are substantially similar. It is a strong signal, not a command: Google weighs it alongside other signals and can select a different canonical than the one declared. Having duplicate URLs from filtering, sorting, tracking parameters, or protocol variants is not automatically a spam problem, it is just something a canonical is meant to sort out.
A canonical is also not a noindex in disguise. One states a preference among duplicates, the other removes an indexed page from results once processed. Confusing the two is one of the more common mistakes in an audit.
The concrete check: compare the canonical the template declares, what the page actually contains, and the canonical Google Search Console reports it selected. Look for templates that collapse genuinely distinct articles onto one generic URL, and make sure the canonical tag actually lives inside a valid head. On a multilingual site, keep this separate from hreflang: locale alternatives need their own coherent set of canonical and alternate annotations, not one instruction that folds every locale into a single URL.
Image alt attribute
The alt attribute lives on the image element itself, as in <img src="[file]" alt="Consultant reviewing a client performance report">. It is not a standalone meta tag, and it is not a hidden keyword field either, even though plenty of old-school seo html elements advice still treats it that way. Its first job is giving someone who cannot see the image a useful description. Google also reads it to understand what the image shows and how it relates to the page around it.
Write what is actually informative in the context of that page. Skip the habit of cramming a list of target phrases into alt text, and skip the copy-pasted boilerplate that gets reused across a whole product gallery, since neither helps a reader or a crawler. A purely decorative image does not need invented, keyword-heavy description bolted on for SEO reasons that do not exist.
Prioritize the images that actually carry information: product shots, diagrams, charts, anything a reader would lose something by not seeing. Do not tell a client every image needs the page's target keyword stuffed into its alt text. And check that lazy-loaded images are actually discoverable when the page renders. An image the crawler never encounters does not benefit from good alt text, because it never gets read.
Structured data markup
Structured data can live in the page as a JSON-LD script, or through in-page formats like Microdata and RDFa. Google recommends JSON-LD where the site setup allows it. What it actually does is describe specific content on the page well enough that an eligible page might become available for a supported rich result. Valid markup does not guarantee that result shows up. Google's general policy is that the markup has to represent what is visible on the page and follow the rules for that particular feature.
Where this goes wrong is scope creep: adding schema types that do not match what a page actually is, marking up claims a reader cannot see, or promising a client a ranking bump because a page now "has schema." None of that is how it works.
The check is a two-step one. First, decide whether the page genuinely matches a Google-supported feature type. Then test the rendered page against that feature's documentation and Google's Rich Results Test, and keep eligibility and actual errors as separate line items in the audit, since a page can be eligible and still throw an error. On a general reference like this one, product and article markup are the two types worth knowing by name; a full schema-type roundup belongs in its own piece.
Hreflang link elements
An HTML link element carrying rel="alternate" and a hreflang value points to the corresponding localized version of a page, so a UK-English page can name its US-English counterpart, and x-default can mark a fallback for visitors whose language or region is not covered by any listed alternative. Google also accepts non-HTML ways of implementing this, so do not insist a client's hreflang has to live in the head if their setup handles it another supported way.
Alternate versions need to reference each other. Missing return links are one of the most common reasons an hreflang cluster breaks, so checking reciprocity is worth more time than checking any single tag in isolation.
This only earns a spot on the checklist for clients with actual alternate-language or regional URLs. Do not map every article to a homepage, and do not treat hreflang as a substitute for translating the page or a shortcut to ranking a single-language site internationally. The HTML lang attribute still matters for accessibility, but Google says it detects a page's language on its own rather than leaning on that attribute to do it.
og:image metadata
Open Graph tags get treated as pure social-sharing metadata, which makes it tempting to call them irrelevant to search. That call is a little too broad. Google's documentation update in March 2026 named both schema.org markup and the og:image tag as sources it can use when choosing preferred image thumbnails in Search and Discover. That is about which image gets shown, not a documented lift to ordinary rankings, and it is worth being precise with a client about that difference.
Where a page has a lead image that matters, check the rendered og:image value and confirm the asset it points to is actually usable. Be clear that this shapes Google's preference for a thumbnail, not a guarantee of which one it picks. And keep this in proportion: a missing og:image should not jump ahead of an accidental noindex or a broken canonical in an audit backlog, because it is a smaller fix with a smaller ceiling.
Valid head and semantic HTML
Google treats the head as the primary place for page metadata, and an invalid element placed inside it can make Google stop reading the head early and ignore whatever metadata comes after. Google specifically calls out misplaced iframe and img elements as examples of what breaks this; the head is meant to hold elements like meta, link, script, and style. That makes basic validity operationally important, even though "valid HTML" itself carries no numeric ranking weight. Check the rendered HTML, not just the template, especially on a site where JavaScript changes metadata after load.
Semantic elements and clear text structure make content easier for both people and crawlers to understand. Google's own guidance points toward semantic HTML and content that is actually present in the DOM, rather than content pulled in through a method the crawler may not process. Use <article> and <main> because they describe what is actually there, not because someone assigned them a point value. The same goes for <strong> and <em>: they should mark real emphasis, not act as a place to sprinkle target phrases.
The move here is to fix broken head output and missing or inaccessible body content before spending more time polishing lower-value metadata. A page with a broken head is losing more than a page with imperfect alt text, and this is usually the seo html elements check that gets skipped because nothing about it looks broken from the outside.
Meta keywords and other dead weight
The meta keywords tag, <meta name="keywords" content="...">, is the one that will not die in client briefs. Google Search says plainly that it has no effect on indexing or ranking. There is no ideal count of phrases, no comma-separated formula, and no version of this tag that works. This is different from words that actually appear in your content or in a structured-data property, those still do something; the meta keywords tag itself does not.
The same goes for revisit-after, an old tag some CMS templates still emit that was never used to tell Google when to recrawl a page. If a client's CMS is also outputting author, generator, or site-verification metadata in the head, do not assume it carries ranking weight just because it sits there. Site-verification metadata has a real job, confirming ownership in Search Console, it just is not an SEO signal.
This is where you save a client time rather than spend it. Delete "fill in meta keywords" from the recurring content brief. Removing an existing meta keywords tag is housekeeping, not a ranking move, so do not sell it to a client as one either way.
What this means for AI search
Google says its existing search fundamentals carry over to AI Overviews and AI Mode, and there are no extra tags or special optimizations that exist purely to get picked up there. The same preview controls covered above, nosnippet, data-nosnippet, max-snippet, and noindex, are what Google names as the way to limit what shows from a page in its AI-driven search features. That is a real reason this list still matters in 2026.
It is not evidence that adding schema, alt text, or an Open Graph tag causes an AI citation on its own, and it does not tell you how ChatGPT, Perplexity, or any other answer engine behaves, since each handles retrieval differently. There is no reliable published percentage lift or tag-weight score behind any item on this list, for Google's own search results or for any other engine, so do not let a client brief invent one.
What to fix first
Not every one of these html tags for seo carries the same weight, so work this list by consequence, not by how old or unfamiliar a tag feels. Start with whether the pages that should be indexed actually are, free of an accidental noindex or an overly tight preview directive. Then check that duplicate URL families point consistently to the canonical you actually want. After that: alt text on the images that carry real information, structured data accuracy on any page eligible for a feature, hreflang coherence on multilingual sites, and finally the hygiene layer, a valid head and semantic structure that keeps everything else readable. Meta keywords and revisit-after do not belong on this list at all; take them off the brief and move on.
Getting this right by hand, across every client and every page, is exactly the kind of mechanical checking that eats an agency's margin. It is also why this markup and metadata gets built into a page at the point it is written rather than patched on afterward in DeepSmith's own production pipeline, so a strategist is reviewing the judgment calls instead of chasing missing alt text across a back catalog.



