You open an AI draft. It reads fine. You still cannot say whether it is good enough to publish, and neither can the person reviewing it after you. That gap is what an AI content quality rubric closes. By the end of this guide you will have a weighted content quality scorecard, a set of pass-or-fail gates, and a threshold that turns "this feels okay" into an approve-or-revise call your whole team makes the same way.
Here is the short answer, so you have it before the detail.
To evaluate AI draft quality objectively, define the draft's job, freeze the version you are reviewing, check its material claims against real evidence, rate seven fixed criteria from 0 to 4, apply weights you set in advance, run your non-negotiable gates, and compare the total to a published threshold. A weighted score gives you consistency. It never outranks a material factual error, an unsupported product claim, or a missing must-have.
The starting policy in this guide: seven criteria, 100 points, publish-ready at 85 or higher only when every gate passes. Those numbers are an internal starting setting, not an industry standard. You will calibrate them on your own drafts.
The rubric at a glance
Seven criteria, 100 weight points, one integer rating each.
| Criterion | Weight | What you are measuring |
|---|---|---|
| Task fit and audience usefulness | 20 | Whether the article does the job in the brief and leaves the reader able to act |
| Factual accuracy and completeness | 25 | Whether claims, numbers, names, dates, quotes, instructions and caveats are correct |
| Evidence, sourcing, and claim integrity | 15 | Whether important claims trace to appropriate, current sources that support the exact wording |
| Brand, product accuracy, and voice | 15 | Whether the draft is right about your company and products, and sounds like you |
| Originality and substantive value | 10 | Whether it adds analysis, examples or synthesis instead of rewriting other pages |
| Clarity, structure, and accessibility | 10 | Whether a real reader can scan, understand and use the page |
| Search and answer readiness | 5 | Whether the page is organized for search and AI answers without going search-first |
| Total | 100 |
Accuracy carries the most weight on purpose. Fluent AI output can be completely coherent and still be wrong, and a wrong claim costs you more trust than a dull paragraph ever will.
Every criterion uses the same 0 to 4 scale, whole numbers only.
| Rating | Meaning |
|---|---|
| 4 | Publish-ready on this criterion. Clean, well evidenced, no material defect. |
| 3 | Meets the requirement. Minor imperfections, nothing that needs correcting. |
| 2 | Mixed or incomplete. A material gap or evidence problem that needs revision. |
| 1 | Major failure. Cannot pass without substantial revision. |
| 0 | Absent, false, unusable or unsafe. A zero on a critical criterion is an automatic fail. |
No half points. Half points feel precise and are not, and they let a reviewer avoid the decision the anchors are asking for.
Step 1: Define what publish-ready means for this draft
Write the standard down before you read a word of the draft.
Do: Freeze the brief. Note the intended reader, the reader's task, the outcome, the target action, the must-have sections, the claims you are allowed to make, and any high-risk subject matter. Decide what the score covers. For a normal web article, score the whole package: title, intro, headings, body, citations, links, metadata, and the accessibility basics.
Done means: Someone who did not write the draft can state what success looks like in two sentences and list the required elements without asking the author.
Where teams go wrong: They change the standard after seeing the draft. A writer cannot fail a requirement nobody wrote down, and you cannot add a preference just because the piece went somewhere you did not expect.
If this feels like extra admin, it is the step that saves you the most argument later. Ten minutes here removes the "well, I would have done it differently" conversation entirely.
Step 2: Freeze the version and build a claim-and-evidence record
Do: Score one saved version, not a document someone is still editing. Then list the material claims before you rate anything. Material claims are the ones worth checking: numbers, dates, prices, limits, names, products, features, quotes, study findings, comparisons, superlatives, product capabilities, anything legal or medical or financial, and any instruction that could push a reader into a consequential action.
Keep a claim ledger with one row per material claim, not one row per paragraph. The fields that matter: claim text, claim type, risk level, source, the exact supporting passage, date checked, status, and action.
Done means: Every material claim is marked supported, corrected, removed, or unresolved. You can also say which required elements are present and which are missing.
Where teams go wrong: They fact-check the intro, or only the lines that sound suspicious. AI errors are rarely suspicious. They are plausible, specific, and sitting in the middle of a paragraph that reads beautifully.
Common mistake: A fluent draft can still carry a fabricated fact or a citation that does not support the sentence attached to it. Check the material claims before you reward the polish.
This is also where stored brand context earns its keep. If your product facts, personas, voice, content types and trusted sources live somewhere structured instead of in a slide deck nobody opens, you can check a draft against them in minutes. That is what Deep IQ does inside DeepSmith: it holds company positioning, products, buyer personas, brand voice, content types and visual guidance as records the writing pipeline is grounded in. Useful for the brand criterion, and useful for the ledger. It is not a factuality certificate. You still test the claims against evidence.
Step 3: Run the non-negotiable gates before you calculate anything
Do: Check the gates first. A weighted total is a summary, not permission to trade a critical standard away.
Six gates to start with:
- Material truth gate. No unresolved factual error, fabricated source, fabricated quote, wrong number, wrong date.
- Claims gate. Objective product or service claims have a reasonable basis before publication. If the wording implies tests, studies, experts or customer results, you need support at that level.
- Brand and product gate. No invented feature, integration, price, limit, customer result, ranking or guarantee. Product names and documented limits are correct.
- Brief gate. Every must-have element is present. A missing deliverable is a task-fit failure even when the writing is excellent.
- Reader-usefulness gate. The page gives a usable answer and does not swap help for search manipulation.
- High-stakes gate. Health, finance, safety, law, security or regulated claims go through your qualified subject-matter or compliance review before publishing.
Done means: Each gate is marked Pass or Fail with one short reason. A failed gate blocks the publish-ready label until it is resolved.
Where teams go wrong: They let a 92-point total hide a false product claim. Gates stop the score turning into a compensation game where good style buys off weak truth.
One gate not to use: an AI detector result. Whether a tool helped write the draft tells you nothing about whether it is accurate, useful, original or compliant. Score the output, not its origin story.
Step 4: Score each criterion independently from 0 to 4
This is where you evaluate AI draft quality criterion by criterion, and the order you do it in matters.
Do: Give one integer to each of the seven criteria. Read the criterion's question, look at the evidence, then write the reason. Score accuracy and evidence before voice and polish, so good prose does not colour the rest.
Here is what each criterion is actually asking.
Task fit and audience usefulness. Does this fulfill the brief for the intended reader, or would that reader have to go searching again? A 4 answers the question directly and gives enough to act on. A 2 is relevant but incomplete, generic, or aimed at the wrong buyer stage. Score it against the brief, never against your own taste.
Factual accuracy and completeness. Can you verify every material statement, with the qualifications that keep it accurate? A 4 has all audited claims accurate, wording that matches the evidence, and no invented facts, numbers, citations or entities. A 2 has at least one claim that is unsupported, overstated, outdated or missing a needed caveat. Completeness is not length. There is no word count that proves an article is good.
Evidence, sourcing, and claim integrity. Can a second reviewer trace the important claims to a source that supports the exact wording? A statement can be true and still poorly supported, so keep this separate from accuracy. Prefer first-party documents, official standards and original studies when the claim is consequential. A reputable secondary summary adds context, it does not replace the original.
Brand, product accuracy, and voice. Does the piece sound like you and describe your products, audience and limits correctly? A 4 gets names, capabilities, value propositions and limitations right, in your voice, with examples that could not be swapped into any other company's blog. A 1 invents or conflates capabilities.
Originality and substantive value. Does it add analysis, examples, experience or synthesis beyond a rewrite of pages that already exist? Original does not mean novel for its own sake. A clear synthesis with a practical framework and audience-specific examples is original enough.
Clarity, structure, and accessibility. Can the reader scan, understand and use the page? Check that headings describe the content rather than decorate it, that heading structure matches the information structure, that the page language is identifiable, and that link purpose is clear from the link text or its immediate context. Resist setting an arbitrary reading-grade or sentence-length cutoff unless you have validated it for your audience.
Search and answer readiness. Is the page easy for a person, a search engine and an AI answer system to understand, without becoming search-first? A 4 has an accurate title, the answer near the top, headings that mirror the questions or steps, natural terminology, accurate metadata and descriptive anchors. A 1 is padded to a length target or overpromises in the title.
Done means: Seven ratings, seven one-sentence reasons, and enough evidence for another reviewer to reproduce your decision. "Sounds good" is not a reason. "3, all steps are present and actionable, but the example is generic" is.
Where teams go wrong: Scoring by overall impression, or punishing the same flaw under all seven headings. One unsupported claim belongs in accuracy and evidence, plus a gate failure if it is material. It does not get to cost you points in clarity too.
If your team runs a production platform, the draft you score is whatever that platform hands you. In DeepSmith, Content Studio produces the researched, brand-grounded article with links, cover image and publish-ready metadata, and it lands in Produced Content for review. The lesson holds either way: a production system can take the repetitive preparation off your plate, and your scorecard still decides whether what came out meets your bar.
Step 5: Calculate the AI content quality score
Do: Convert each rating with one formula.
Weighted contribution = criterion weight × rating ÷ 4
AI content quality score = the sum of all seven contributions
The maximum is 100. Weights are fixed before scoring. If a criterion genuinely does not apply, mark it N/A before you read the draft and use a normalization rule you declared in advance. Never drop a criterion after reading, because that is just deleting the part the draft failed.
An example of the arithmetic:
| Criterion | Weight | Rating | Contribution |
|---|---|---|---|
| Task fit and audience usefulness | 20 | 4 | 20.00 |
| Factual accuracy and completeness | 25 | 3 | 18.75 |
| Evidence, sourcing, and claim integrity | 15 | 3 | 11.25 |
| Brand, product accuracy, and voice | 15 | 4 | 15.00 |
| Originality and substantive value | 10 | 3 | 7.50 |
| Clarity, structure, and accessibility | 10 | 4 | 10.00 |
| Search and answer readiness | 5 | 3 | 3.75 |
| Total | 100 | 86.25 |
That draft lands on 86.25 with nothing below a 3, so it would clear the starting threshold if the gates passed. It is an illustration of the maths, nothing more. It is not a benchmark and not anyone's result.
One thing worth saying out loud to your team: 86.25 does not mean the article is 86 percent true. It is a weighted decision score across seven different dimensions. Truth is checked claim by claim, and a single material error fails a gate no matter how high the total climbs.
Done means: The score is reproducible from the weights and the ratings, and the label agrees with the gate results.
Where teams go wrong: Reporting a total with no breakdown. The number alone cannot tell the writer what to fix, and it hides whether a high score came from one strong dimension carrying a weak one.
Step 6: Calibrate the scorecard across reviewers
An AI content quality rubric only creates consistency once two people use it the same way. That takes one deliberate session, not a quarter.
Do: Give the same saved drafts to at least two reviewers. Require independent ratings before anyone talks. Use a mix from your own content: clear approvals, clear revisions, and the borderline ones that started this whole problem. Compare at criterion level, not just totals.
Rules that keep the session honest:
- Everyone uses the same version, brief, evidence pack, weights and anchors.
- A gap of more than one point on any criterion triggers an evidence-based discussion.
- Any disagreement on a gate blocks approval until it is settled.
- Never average away a gate failure.
- When the same disagreement keeps returning, sharpen the anchor or add a worked example. Do not quietly lower a weight so the draft passes.
- Keep a small library of scored examples and update it when your interpretation shifts.
Done means: Reviewers can explain their ratings, and repeat disagreements get fixed by better definitions rather than by whoever is most senior.
Where teams go wrong: Asking an automated grader to set the standard. Automated or LLM-based scoring can help with a pre-score, a filter or a regression check, and it should use the same detailed rubric and be validated against human labels first. The publication call stays with a person.
Pro tip: Whole-number ratings plus one evidence sentence per score. A score with no reason cannot be calibrated, which means it cannot be improved.
Step 7: Make the approve-or-revise call and record why
Do: Apply the threshold only after the gates and the seven ratings are done.
Starting decision bands:
- 85 to 100, all critical gates passed: publish-ready.
- 70 to 84, or any gate failed: revise, then score the revised version again.
- 0 to 69: not publish-ready. Send it back for substantive revision and re-score.
Plus these minimums for a publish-ready label: at least 3 on task fit, at least 3 on accuracy, at least 3 on evidence, at least 3 on brand and product accuracy, and no criterion rated 0.
Done means: The record holds the version, reviewer, date, seven ratings, weighted total, gate results, unresolved issues and the final label. Six months from now someone can see exactly why this went live.
Where teams go wrong: Treating 85 as universal truth. It is a practical starting point. Once you have enough scored examples, adjust weights or thresholds on purpose, write down what changed, and keep the old version so your historical comparisons still mean something.
This is also the moment the phrase publish-ready earns a definition. It describes a draft that cleared your standard, not a draft that came out of a tool. In DeepSmith, Produced Content is where you review, edit and publish, and that sequence makes the same point: the tool prepares the article, your scorecard decides.

What to do next
Save the content quality scorecard somewhere your reviewers actually work. Write your own anchors for each criterion, using one approved example and one revised example from your own content. Run a calibration session on five real drafts. Then score every piece the same way for a month and see whether the borderline arguments get shorter.
You do not need a perfect version to start. You need one written-down version that two people can use to score AI draft quality the same way this week.
If you want to see what a brand-grounded draft looks like before you commit, start a free DeepSmith trial and run one of your own topics through it. Then score it with the rubric you just built. That is the honest test.



