DeepSmith

Sep 26 · AEO & AI Visibility

14 min read

Which On-Page HTML Tags Still Matter for SEO (and Which Are Dead Weight)

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
Abstract monochrome illustration of HTML tag brackets connected in a network, some linked and glowing, some faded and disconnected, with the cover line HTML Tags That Still Matter centered over them.

You already know title tags, headings, and meta descriptions matter. Those get their own deep dives elsewhere. This is the list for everything else: the smaller pieces of markup that show up in every client audit, the ones that get argued about in briefs and Slack threads without anyone checking whether Google still cares. Some of these html tags for seo still do real work. A few never did, and one or two are pure legacy weight that somebody copied into a template years ago and nobody has removed since.

"Matters for SEO" does not mean "is a ranking factor" here, and that distinction is the whole point of this list. A robots directive can pull a page out of results entirely. A canonical tag can shift which URL Google shows for a set of duplicates. Structured data can make a page eligible for a search feature it would not otherwise get. None of that is the same as a keyword boost, and treating it that way is how audits end up wasting a client's dev hours on the wrong fix.

Here is the fast reference, then each item with what it actually does and the move to make on a real site. Think of it as the shortlist of seo html elements worth an agency's time once the bigger pieces are handled.

MarkupVerdictWhat it actually does
Robots meta directivesHigh priority when indexing or previews need controlTells a crawler whether a page may be indexed and how it may be shown
Canonical link elementHigh priority wherever duplicate URLs existStates a preferred URL among near-duplicates; Google can still choose differently
Image alt attributeWorth doing on meaningful imagesDescribes an image for accessibility and helps Google understand it
Structured data markupConditional, valuable for eligible page typesDescribes page content and can qualify it for a search feature, not guaranteed to appear
Hreflang link elementsEssential only for real locale alternativesPoints to the corresponding version of a page in another language or region
og:image metadataConditional display control, not a ranking leverCan inform which thumbnail Google shows in Search and Discover
Valid head and semantic HTMLFoundational hygieneKeeps metadata readable and content accessible to crawlers
Meta keywords and old scheduling tagsDrop from the checklistDo nothing for Google indexing or ranking

Robots meta directives

A page-level tag like <meta name="robots" content="noindex"> sits in the document head and tells Google not to show that page in results, once Google has actually fetched and processed the directive. That last part trips people up constantly. If the URL is blocked in robots.txt, Google may never see the on-page noindex at all, so the two controls are not interchangeable and blocking a page in robots.txt is not the same as removing it from the index.

The defaults are index and follow, so writing index,follow on every page does nothing useful. Where this tag earns its place is preview control: nosnippet, max-snippet, max-image-preview, and the text-level data-nosnippet attribute all shape what Google may show from a page without touching whether the page ranks. A googlebot-specific version of the tag addresses Google only, and when directives conflict, the more restrictive one wins.

The agency move here is simple and worth doing before anything else on this list: check the rendered head of a page the client believes is indexable, and confirm there is no accidental noindex or an overly restrictive snippet directive left over from a staging build. This is also the family Google points to when a publisher wants to limit what shows up in its AI-driven search features, so it is worth checking on pages a client wants kept out of AI Overviews. It does not extend to controlling every other AI engine, and you should not tell a client it does.

The canonical link element, <link rel="canonical" href="[preferred URL]">, states which URL you would rather Google treat as the representative one when several pages are substantially similar. It is a strong signal, not a command: Google weighs it alongside other signals and can select a different canonical than the one declared. Having duplicate URLs from filtering, sorting, tracking parameters, or protocol variants is not automatically a spam problem, it is just something a canonical is meant to sort out.

A canonical is also not a noindex in disguise. One states a preference among duplicates, the other removes an indexed page from results once processed. Confusing the two is one of the more common mistakes in an audit.

The concrete check: compare the canonical the template declares, what the page actually contains, and the canonical Google Search Console reports it selected. Look for templates that collapse genuinely distinct articles onto one generic URL, and make sure the canonical tag actually lives inside a valid head. On a multilingual site, keep this separate from hreflang: locale alternatives need their own coherent set of canonical and alternate annotations, not one instruction that folds every locale into a single URL.

Image alt attribute

The alt attribute lives on the image element itself, as in <img src="[file]" alt="Consultant reviewing a client performance report">. It is not a standalone meta tag, and it is not a hidden keyword field either, even though plenty of old-school seo html elements advice still treats it that way. Its first job is giving someone who cannot see the image a useful description. Google also reads it to understand what the image shows and how it relates to the page around it.

Write what is actually informative in the context of that page. Skip the habit of cramming a list of target phrases into alt text, and skip the copy-pasted boilerplate that gets reused across a whole product gallery, since neither helps a reader or a crawler. A purely decorative image does not need invented, keyword-heavy description bolted on for SEO reasons that do not exist.

Prioritize the images that actually carry information: product shots, diagrams, charts, anything a reader would lose something by not seeing. Do not tell a client every image needs the page's target keyword stuffed into its alt text. And check that lazy-loaded images are actually discoverable when the page renders. An image the crawler never encounters does not benefit from good alt text, because it never gets read.

Structured data markup

Structured data can live in the page as a JSON-LD script, or through in-page formats like Microdata and RDFa. Google recommends JSON-LD where the site setup allows it. What it actually does is describe specific content on the page well enough that an eligible page might become available for a supported rich result. Valid markup does not guarantee that result shows up. Google's general policy is that the markup has to represent what is visible on the page and follow the rules for that particular feature.

Where this goes wrong is scope creep: adding schema types that do not match what a page actually is, marking up claims a reader cannot see, or promising a client a ranking bump because a page now "has schema." None of that is how it works.

The check is a two-step one. First, decide whether the page genuinely matches a Google-supported feature type. Then test the rendered page against that feature's documentation and Google's Rich Results Test, and keep eligibility and actual errors as separate line items in the audit, since a page can be eligible and still throw an error. On a general reference like this one, product and article markup are the two types worth knowing by name; a full schema-type roundup belongs in its own piece.

An HTML link element carrying rel="alternate" and a hreflang value points to the corresponding localized version of a page, so a UK-English page can name its US-English counterpart, and x-default can mark a fallback for visitors whose language or region is not covered by any listed alternative. Google also accepts non-HTML ways of implementing this, so do not insist a client's hreflang has to live in the head if their setup handles it another supported way.

Alternate versions need to reference each other. Missing return links are one of the most common reasons an hreflang cluster breaks, so checking reciprocity is worth more time than checking any single tag in isolation.

This only earns a spot on the checklist for clients with actual alternate-language or regional URLs. Do not map every article to a homepage, and do not treat hreflang as a substitute for translating the page or a shortcut to ranking a single-language site internationally. The HTML lang attribute still matters for accessibility, but Google says it detects a page's language on its own rather than leaning on that attribute to do it.

og:image metadata

Open Graph tags get treated as pure social-sharing metadata, which makes it tempting to call them irrelevant to search. That call is a little too broad. Google's documentation update in March 2026 named both schema.org markup and the og:image tag as sources it can use when choosing preferred image thumbnails in Search and Discover. That is about which image gets shown, not a documented lift to ordinary rankings, and it is worth being precise with a client about that difference.

Where a page has a lead image that matters, check the rendered og:image value and confirm the asset it points to is actually usable. Be clear that this shapes Google's preference for a thumbnail, not a guarantee of which one it picks. And keep this in proportion: a missing og:image should not jump ahead of an accidental noindex or a broken canonical in an audit backlog, because it is a smaller fix with a smaller ceiling.

Valid head and semantic HTML

Google treats the head as the primary place for page metadata, and an invalid element placed inside it can make Google stop reading the head early and ignore whatever metadata comes after. Google specifically calls out misplaced iframe and img elements as examples of what breaks this; the head is meant to hold elements like meta, link, script, and style. That makes basic validity operationally important, even though "valid HTML" itself carries no numeric ranking weight. Check the rendered HTML, not just the template, especially on a site where JavaScript changes metadata after load.

Semantic elements and clear text structure make content easier for both people and crawlers to understand. Google's own guidance points toward semantic HTML and content that is actually present in the DOM, rather than content pulled in through a method the crawler may not process. Use <article> and <main> because they describe what is actually there, not because someone assigned them a point value. The same goes for <strong> and <em>: they should mark real emphasis, not act as a place to sprinkle target phrases.

The move here is to fix broken head output and missing or inaccessible body content before spending more time polishing lower-value metadata. A page with a broken head is losing more than a page with imperfect alt text, and this is usually the seo html elements check that gets skipped because nothing about it looks broken from the outside.

Meta keywords and other dead weight

The meta keywords tag, <meta name="keywords" content="...">, is the one that will not die in client briefs. Google Search says plainly that it has no effect on indexing or ranking. There is no ideal count of phrases, no comma-separated formula, and no version of this tag that works. This is different from words that actually appear in your content or in a structured-data property, those still do something; the meta keywords tag itself does not.

The same goes for revisit-after, an old tag some CMS templates still emit that was never used to tell Google when to recrawl a page. If a client's CMS is also outputting author, generator, or site-verification metadata in the head, do not assume it carries ranking weight just because it sits there. Site-verification metadata has a real job, confirming ownership in Search Console, it just is not an SEO signal.

This is where you save a client time rather than spend it. Delete "fill in meta keywords" from the recurring content brief. Removing an existing meta keywords tag is housekeeping, not a ranking move, so do not sell it to a client as one either way.

Google says its existing search fundamentals carry over to AI Overviews and AI Mode, and there are no extra tags or special optimizations that exist purely to get picked up there. The same preview controls covered above, nosnippet, data-nosnippet, max-snippet, and noindex, are what Google names as the way to limit what shows from a page in its AI-driven search features. That is a real reason this list still matters in 2026.

It is not evidence that adding schema, alt text, or an Open Graph tag causes an AI citation on its own, and it does not tell you how ChatGPT, Perplexity, or any other answer engine behaves, since each handles retrieval differently. There is no reliable published percentage lift or tag-weight score behind any item on this list, for Google's own search results or for any other engine, so do not let a client brief invent one.

What to fix first

Not every one of these html tags for seo carries the same weight, so work this list by consequence, not by how old or unfamiliar a tag feels. Start with whether the pages that should be indexed actually are, free of an accidental noindex or an overly tight preview directive. Then check that duplicate URL families point consistently to the canonical you actually want. After that: alt text on the images that carry real information, structured data accuracy on any page eligible for a feature, hreflang coherence on multilingual sites, and finally the hygiene layer, a valid head and semantic structure that keeps everything else readable. Meta keywords and revisit-after do not belong on this list at all; take them off the brief and move on.

Getting this right by hand, across every client and every page, is exactly the kind of mechanical checking that eats an agency's margin. It is also why this markup and metadata gets built into a page at the point it is written rather than patched on afterward in DeepSmith's own production pipeline, so a strategist is reviewing the judgment calls instead of chasing missing alt text across a back catalog.

Frequently asked questions

Do meta keywords help SEO in 2026?

Not in Google Search. Google says the meta keywords tag does not affect indexing or ranking, so effort spent maintaining it is effort not spent on content or the markup that actually has a job.

Is image alt text a ranking factor?

Google uses it to understand an image and its context, and it has a real accessibility purpose, but there is no documented ranking weight attached to it. Write it for the reader and the crawler, not to hit a keyword quota.

Are canonical tags and robots directives interchangeable?

No. A canonical states a preferred URL among duplicates, and Google can still choose a different one. A processed noindex removes a page from results outright. Confusing the two is a common and avoidable audit mistake.

Does structured data guarantee a rich result or an AI citation?

No. Accurate markup can make an eligible page qualify for a supported search feature, but qualifying and actually appearing are two different things, and the same is true for AI citation.