DeepSmith

Sep 26 · Content Operations

19 min read

How to Run a 90-Day Content Pilot for a New Client and Prove It Works

Avinash Saurabh
Avinash Saurabh · CO-Founder & CEO
A monochrome illustration showing measurement tracks running from a baseline flag on the left, past evenly spaced checkpoint markers and faint chart bars, to a checkmark badge at the right, under the words Prove It in One Quarter.

A new client has signed a small first engagement. They want to see something real before they commit to a year of retainer. That is fair, and it is also a test of you.

This guide walks you through a 90 day content pilot you can actually defend at the end. You will leave with a locked scope, a saved baseline, a small scorecard, three phases of work, and a final review that ends in a clear decision. Feeling behind before you start? That is normal. The first step is smaller than you think.

Be honest about one thing up front. A quarter can prove you have a credible process, that the two teams can work together, that the agreed work shipped, and that early signals moved. It cannot prove content alone caused revenue, or that every page will rank. Say that on day one and you protect yourself on day 90.

Step 1: Name the one decision the pilot has to settle

Every content pilot for new client work starts with a single question. Not a goal, a decision.

Good ones sound like this:

  • Should this client retain us for an ongoing content program?
  • Can we produce work that reaches their priority audience and meets their editorial standard?
  • Can the agreed themes generate early search, engagement, AI-visibility, or qualified-lead signals?
  • Does our workflow produce enough evidence to justify expanding?

Then turn the client's broad goal into something specific. "Increase awareness" and "generate more leads" are too vague to govern a quarter. A useful objective names the change you expect and how it connects to the business.

Now write the content pilot plan down. One page, agreed by both sides:

  • Start and end dates.
  • The target audience or buyer segment.
  • The content type and the topic boundary.
  • How many pieces, if you both agree on a number.
  • Whether the work is new content, updates to existing pages, or both.
  • Who publishes.
  • Who reviews, and who has final approval authority.
  • Which systems you need access to: analytics, Search Console, CMS, CRM.
  • Where the measurements come from.
  • The review dates.
  • The final decision: continue, revise, extend, or stop.
  • What is explicitly excluded. A site migration. A full technical SEO program. Paid media. Unlimited content.

That last bullet saves more agency content trial engagements than any other line on the page. Write the exclusions before someone asks for them.

A content pilot plan this specific also does quiet sales work. It shows a client you have run this before.

Done when: you and the client can answer seven questions on one page. What are we testing? For whom? With what work? By when? Against which baseline? Which signals count as progress? What would make us continue, change, or stop?

Common mistake: promising a full marketing turnaround in a quarter. A pilot tests process, fit, execution, and early evidence. It is not a guarantee of mature organic growth, and selling it as one is how agencies lose renewals they earned.

There is no standard article count that makes a pilot valid. Set the volume from the client's resources, your capacity, how hard the topics are, how fast approvals move, and how much work it takes to make the test mean something.

Step 2: Get access, context, and one named decision-maker

Month one is not a race to publish. It is the month you build the foundation everything else stands on.

Ask for access to what the agreed measurement and publishing work needs: analytics, Search Console, the CMS, the CRM, search or ad platforms if those are in scope, the existing content inventory, existing reporting, and the approval and legal review process.

Then get the business context that no login gives you:

  • What has worked before, and what has failed.
  • Which audience and buyer segment matters most.
  • Which products, services, and claims are strategic.
  • Which competitors matter.
  • Who approves content.
  • Whether legal, compliance, brand, or executive review is required.
  • How fast reviewers actually return feedback.

Appoint one client-side point person. One. They gather feedback, route requests, and keep approvals moving. Document the approval path and the longest review time you expect. If a CEO or a legal team has to see every piece, that dependency belongs in the plan, not in a surprise on day 60.

Treat the first weeks as a listening tour. Interview marketing leadership, sales, support, product, and the people who talk to customers. Read sales and support calls. Ask what buyers object to.

Then turn all of that into one source of truth: positioning, audience and pain points, the company description and differentiators, product facts, approved and prohibited claims, the writing voice, reference content showing the expected standard, and the buyer questions you plan to answer.

This is where a multi-client roster gets expensive. Every account has its own voice, its own product facts, and its own no-go claims, and holding them all in people's heads is what makes onboarding take weeks. DeepSmith stores that context as structured records in Deep IQ, so each client's positioning, personas, product facts, and voice sit in one place that every draft is grounded in. Each client also gets its own workspace, with its own brand, content, and reporting from a single account, so one client's voice never leaks into another's draft.

Done when: you have access, a named decision-maker, an approval map, an approved brief, a documented voice, and enough context to write the first piece without guessing.

Where pilots go wrong: the real risks here are rarely writing risks. Expectation gaps. Thin onboarding. Unclear quality standards. Legal delays. Late executive revisions. The senior team disappearing after the pitch. Treat each one as a pilot variable you record, not a problem you hide until the final review.

Step 3: Capture the baseline before you publish anything

You cannot prove movement without a starting point. Save the baseline before the work changes a single page.

Record the date range, the filters, the dimensions, the URLs, the queries, and the exports, so another analyst could reproduce your final comparison. At minimum, capture:

  • How the target pages perform today.
  • Search clicks, impressions, click-through rate, and average position.
  • The queries those pages already show up for.
  • Organic sessions or users, where analytics allows.
  • The engagement metrics that match your objective.
  • Existing key events or conversions.
  • Leads or qualified leads, where the CRM and analytics support it.
  • The existing content inventory and page status.
  • Backlinks or other agreed authority indicators, if they are in scope.
  • Existing AI-search visibility, if AI visibility is part of the pilot.

Get the definitions right, because clients will ask. Impressions count how often a link to the site was seen or could have been seen in Google Search, Discover, or News, and some result types only count once the item is scrolled into view. Clicks count how often someone clicked through from Google. CTR is clicks divided by impressions. Average position is a relative Google Search position, and it means different things in different situations. It is not a simple rank number, and treating it like one is the fastest way to lose a client's trust.

Keep the setup consistent all quarter. Same property. Same search type. Same date-range logic. Same comparison method. Look at pages and queries for the target content, use weekly or monthly granularity so daily noise does not scare anyone, and export at baseline, mid-pilot, and final review. Treat the most recent data carefully, because it can still change.

One more detail worth knowing: Search Console often assigns data for page variations to the canonical URL Google picked. Make sure your team knows which URL is actually being measured.

If AI visibility is in scope, baseline it the same way you baseline search. DeepSmith tracks how AI engines answer the questions in your client's space and reports mention rate, citation rate, and share of voice, with the answers and cited pages behind each number. Engine coverage is tiered by plan, so check the plan you are on before you promise a client a specific engine.

The AI Visibility overview reports mention rate, citation rate and share of voice as separate top-line metrics, with a per-engine breakdown and a competitor leaderboard showing where the tracked brand ranks against its rivals.

Done when: there is a dated baseline pack another analyst could reproduce. Target page list, baseline metrics, baseline queries, conversion definitions, AI-visibility measurements if relevant, and the exact reporting settings used.

Pro tip: freeze the target-page list on day one. If pages get added later, label them as new entrants instead of quietly folding them into the before-and-after. A comparison that changed its own inputs proves nothing, and a sharp client will spot it.

Step 4: Pick a small scorecard with three layers

Here is the good news: you need fewer metrics than you think. The discipline is choosing the ones that decide something.

Build each measure the same way. Objective, the change the client wants. Key result, the measurable definition of success. KPI, the combined view of progress. Underlying metrics, the individual data points. Analytics method, where and how you collect it.

Then sort your scorecard into three layers.

Layer one, delivery and learning. Access completed. Brief approved. Target pages selected. Pieces produced and approved. Review turnaround. Revision cycles. Instrumentation done. Blockers cleared. These do not prove market value. They prove the experiment was actually run, which matters more than people expect when results are thin.

Layer two, early audience and search signals. Impressions, clicks, CTR, position read cautiously, organic sessions, engagement rate, average engagement time, return visits, shares and link clicks. Where AI search is in scope, add mention rate, citation rate, share of voice, and page-level citation data. These can move inside a quarter, which is why they carry the proof.

Layer three, business outcomes. Qualified leads. Key-event completions. Content-assisted conversions. Lead quality. Opportunity creation. Content consumption by known leads in the CRM. Use these only where the client's systems genuinely support them.

Do not force a revenue conclusion the tracking cannot carry. A pilot can honestly prove better delivery, stronger engagement, improved search visibility, or real qualified demand without proving closed revenue. That is still a win.

Report qualified leads, not total leads. A total lead count looks great and tells a client nothing about whether the right people showed up.

Good sounds like this: "For the agreed target pages, we compare Search Console clicks, impressions, CTR, queries, and average position against the baseline, then report qualified key events and lead quality where tracking exists."

Bad sounds like this: "We will track traffic, rankings, engagement, leads, pipeline, revenue, social, backlinks, and AI visibility," with no word on which numbers decide continuation.

Done when: the scorecard holds a small number of agreed measures, each with a definition, a data source, an owner, a reporting frequency, a baseline, and a rule for how to read it.

If the pilot pushes content through newsletters, social, partners, or campaigns, tag every link the same way from the start. Retrofitting this in week nine does not work.

Google's URL builders support utm_source for the referrer, utm_medium for the marketing medium, utm_campaign for the campaign, utm_content to separate creative or links inside one campaign, utm_term for a paid keyword, and utm_id for a campaign ID. At minimum, use source, medium, and campaign on every tagged link.

Parameter values are case sensitive. Agree on a lowercase standard, one source value per platform and one medium value per channel, and write it down where the whole team can see it.

GA4's Traffic acquisition report is where this pays off. It shows where visitors come from across session campaign, default channel grouping, medium, source, source and medium, and source platform. On the metric side you get sessions, engaged sessions, engagement rate, average engagement time per session, key events, session key-event rate, event count, and revenue.

GA4 calls the business actions that matter "key events," and a key event can also become a conversion in Google Ads. Define yours before the pilot starts, then confirm it is actually firing and reporting. A key event nobody checked is a hole you find at the review.

For a content pilot, connect the source and campaign view to the target page and see whether visitors take the agreed key event. A tagged visit is evidence of a path, not proof that the article caused a sale.

Done when: you can follow a tagged link into the acquisition report, see the target content page, and identify the agreed key event where the setup supports it.

Common mistake: letting LinkedIn, linkedin, and LI all live in the data as different sources. Case-sensitive sprawl fragments a small sample and makes a decent pilot look inconclusive.

Step 6: Run the quarter in three phases

A 90 day content pilot runs in three phases, each with its own job. Do not blur them.

Days 1 to 30: learn and lay the foundation

Kick off and define success in the client's own words. Get analytics, CMS, Search Console, and CRM access. Interview marketing, sales, support, product, and the subject-matter experts. Review past campaigns and existing content. Identify the priority audience, their objections, and their buying stages. Approve the messaging and voice framework. Inventory the existing assets and how they perform. Decide which pages to refresh, expand, keep, or leave out. Select the target set. Capture the baseline. Confirm instrumentation and campaign naming. Agree the reporting cadence.

The deliverable is a short pilot brief: scope, audience, messaging, target content set, baseline, scorecard, owners, approval path, and a risk register.

The client approves that brief and the baseline before the first real batch publishes. If approval slips, record the delay as a pilot dependency. Lost weeks caused by a stalled review are not your underperformance, and the record is what lets you say so calmly at the review.

Failure mode: publishing something flashy in week two, before anyone understands positioning, product facts, audience language, or approvals. It creates rework and weakens the whole test.

Days 31 to 60: build, publish, and inspect the early signals

Produce and get approval on the first batch. Publish or refresh the target content. Apply the agreed internal linking, metadata, SEO, and distribution work. Watch indexing and search performance. Review queries, pages, impressions, clicks, CTR, and position. Check key events and lead quality where you have them. Compare what shipped against the scope. Record client feedback, approval speed, and revision patterns. Then let evidence, not preference, shape the next batch.

This is also where production volume becomes the constraint. Every hour a strategist spends on research, briefs, internal linking, and metadata is an hour that does not scale across a roster. DeepSmith produces publish-ready articles with the research, SEO and AEO structure, internal and external links, and a cover image already built in, each one grounded in that client's stored context. Your strategist reviews for judgment instead of mechanics, which is what makes a pilot deliverable on time without borrowing capacity from your other accounts.

The midpoint review is not a status meeting. It answers real questions. What shipped? What was delayed, and why? Which pages got impressions or clicks? Which queries are emerging? Did engagement or key events move? Are the leads relevant to the agreed persona? What does the data support doing next?

By the end of this phase, both sides should know whether the partnership works operationally. Communication, feedback speed, blocker resolution, work quality. Not just the numbers.

One caveat to repeat here, out loud: Google says some search changes appear in a few days and others take several months. Report movement without promising every page reaches its mature performance by day 90.

Days 61 to 90: run small experiments and prepare the recommendation

Map the buyer journey and find the biggest content gaps. Then pick one or two focused experiments. Test one meaningful variable at a time where you can: a topic, a format, an audience segment, a refresh approach, a distribution route. Keep each one small enough to fail cheaply and specific enough to teach you something.

Ideas worth testing:

  • Refresh a page with impressions but weak CTR.
  • Expand a page that ranks but ignores the buyer's next question.
  • Create a decision-stage page where the inventory shows a gap.
  • Try a different format for the same audience problem.
  • Test a new distribution source using the parameters you agreed in step five.
  • Use AI-visibility data to prioritize a topic a competitor is winning.

Keep measuring the original frozen set separately from the new experimental pages. Mixing them turns a clean comparison into a muddy one.

None of these are guaranteed growth tactics. Their value is partly the evidence they produce, and evidence is what you are selling.

Step 7: Hold the final review and make an honest call

Build the final review as a decision document, not a dashboard tour. Ten parts:

  1. The objective the pilot was meant to test.
  2. The scope, including what was excluded.
  3. The baseline, with starting values and date range.
  4. The delivery record: what shipped, what was approved, what was blocked.
  5. Leading indicators: search visibility, impressions, clicks, CTR, engagement, AI visibility.
  6. Lagging indicators: key events, qualified leads, opportunities, influenced deals, where available.
  7. Experiment results, and what each one taught you.
  8. Attribution limits: what cannot be confidently assigned to the content.
  9. Operational fit: communication, approval speed, quality, collaboration.
  10. Your recommendation.

Give the recommendation one of three shapes. Continue and expand, when the work met the standard, the relationship functions, the measurement is credible, and the signals justify more. Continue with a revised or narrower pilot, when the process works but the test was too short, the content set too small, the tracking incomplete, or the window shortened by approval delays. Stop or pause, when access and approvals never arrived, the work did not meet the standard, or the evidence does not justify more investment.

Do not tie continuation to a percentage threshold unless you both agreed that number before the pilot started. No universal benchmark exists for traffic growth, rankings, leads, citations, or revenue from one quarter of content, and inventing one during an agency content trial is a promise you will be held to.

Then describe your proof at the level the data actually supports. Four rungs, and it helps to name them for the client:

  • Observed: a metric changed in the data source.
  • Associated: the change happened alongside the pilot work.
  • Attributed: the analytics model gave the content some credit.
  • Causally proven: the pilot isolated the effect against a credible counterfactual.

Most 90 day content pilot engagements can support the first three far more easily than the fourth. Calling an association causal is the one shortcut that will cost you the account later.

A rising staircase of four blocks labelled Observed, Associated, Attributed and Causally proven, with a dashed divider showing that most pilots reach the first three levels of proof in one quarter and rarely reach the fourth.

So say the true thing. "The agreed content set produced a measurable change in the selected metric against the recorded baseline, under the documented scope and method." Or: "Content was associated with qualified key events during the pilot, and the acquisition data indicates a contribution. It does not isolate content as the sole cause." Or simply: "We delivered the agreed work, held the measurement system, ran the reviews, and made documented decisions from the data."

What not to say: that the pilot proved content caused all new revenue, that every article will rank in 90 days, that a traffic bump proves the content worked, that no leads means the content failed when the conversion path was never instrumented, or that AI-visibility gains are guaranteed.

What to do next

Turn the review into a written decision. Continue, revise the scope, extend the test, or stop.

Then scale narrowly. Only the audience, the topic, the format, the distribution route, and the workflow the pilot gave the client reason to trust. Not a bigger calendar because the quarter felt good.

You do not need a perfect quarter. You need a documented one. Run it the same way for the next account and the one after, and you stop improvising: you prove content value quarter after quarter with a repeatable method instead of a fresh argument each time. That is what turns a content pilot for new client accounts into a service you can sell again next month.

If AI visibility is part of what you are proving, you can start a free trial and see real tracking data and real drafts for one client account before you pay for anything.

Frequently asked questions

Can a 90-day content pilot prove SEO works?

It can show whether the agreed work was delivered and whether early search, engagement, conversion, or AI-visibility signals moved against a recorded baseline. It cannot guarantee that all SEO effects mature inside one quarter. Google notes that changes may take from a few days to several months to appear, and not every change produces a noticeable result.

How many articles should the pilot include?

There is no universal number. Set the volume from the client's goals, topic difficulty, your capacity, review speed, and publishing resources. A smaller, clearly defined test beats an arbitrary volume nobody can review or measure properly.

Which metrics should the client see at the final review?

The agreed baseline and current values for the target set. Search metrics such as clicks, impressions, CTR, queries, and average position. Then engagement, key events, qualified leads, or content-assisted conversions where the setup supports them. Include the delivery and approval record too, because a test weakened by missing access or slow approvals should not be presented as a pure content-performance result.

Can content be credited with a lead or a sale?

It can be associated with one, and an attribution model may assign it credit. Content journeys are rarely linear, though, so some influence stays directional. Use acquisition data, consistent campaign parameters, key events, CRM content consumption, and the client's attribution model, then state plainly what the data proves and what it does not.