To evaluate a marketing agency, judge the system it will install, not the scope it will bill. Ask what runs after onboarding, who owns the number, how results are reported in revenue, and what happens when a metric falls. Get those four answers in writing before you sign — activity lists and case-study decks won't tell you.

Most owners evaluate agencies the way they'd evaluate a menu: compare deliverables, compare price, pick one. That's how you end up eighteen months later with a dashboard full of impressions, a project manager you've never met, and no clear line between what you paid and what you earned. The diligence below is built for the person who signs the cheque and carries the revenue number personally.


Why do most agency evaluations fail before the contract is signed?

Because they measure the wrong thing. A proposal describes activity — posts, ads, audits, calls. Revenue comes from a system: an offer, a funnel, follow-up, a CRM that actually fires, and tracking that ties a dollar in to a dollar out. An evaluation that never asks about the system can't predict the result.

The accountability gap is well documented. An industry survey distributed via Businesswire found that 71% of brands report frustration demonstrating marketing ROI effectiveness. That number isn't a story about bad marketers. It's a story about engagements structured so no one was ever obligated to prove the connection to money in the first place. If ROI proof isn't built into the agreement, it doesn't appear later by goodwill.

The same pattern shows up in public review data. Reading through G2 and Clutch reviews of agencies (not ours), the recurring complaints aren't about creative taste. They're structural: "constantly changing project managers and the issues with communication made it difficult to work with them"; "the inability to help solve issues... made it so we ended up doing much of the transition ourselves." Owners also describe agencies as "cookie-cutter," with "processes are confusing" and teams lacking the capacity to change direction.

Every one of those failures was visible at evaluation stage — if the right questions had been asked.


What should you ask about the system, not the scope?

Ask what will be running in your business ninety days after kickoff, and who touches it. A scope tells you what an agency will do. A system description tells you what will exist — and whether it keeps producing when the retainer conversation gets awkward.

Use these five questions verbatim:

  1. "What is installed and running at day 90, and can you name each component?" You want an inventory: the offer and message, the funnel, the follow-up sequences, the CRM configuration and automations, the reactivation engine, the acquisition channels, the tracking layer. Vague answers here predict vague results.
  2. "Who owns the outcome, by name, and are they in this room?" If the people pitching disappear after signature, you've bought a hand-off. That's the single most common structural failure in agency work — we cover it in depth in what changes when an agency owns the outcome instead of the scope.
  3. "What number is this engagement accountable to?" Not sessions. Not impressions. A revenue figure, a pipeline figure, or a close-rate figure — something that appears on your P&L. If you're unsure which number is the right one, start with revenue vs profit: which number should your marketing be held to.
  4. "Show me a report you sent a client last month." Redacted is fine. You're looking for whether the report explains movement in a business number or narrates activity. Most reporting proves nothing — here's what separates the two.
  5. "What happens the month a number falls?" The right answer is a diagnostic process, not reassurance. Agencies that explain away a falling number will do it to you for a year.

"They know exactly how to connect marketing execution to real business outcomes." — Riggs Eckleberry, Chairman, OriginClear

That sentence is the whole test, phrased as a compliment. Can this agency draw the line from an action they take to a dollar you receive? If they can't draw it for you during the sale — when they're most motivated — they won't draw it in month seven.


How do you check whether an agency will actually own the result?

Look at how the contract is written. Scope-based agreements define completion as work delivered. Outcome-based agreements define completion as result produced. The difference determines who absorbs the risk when something doesn't work.

Three practical checks:

  • Read the termination and reporting clauses first, not the deliverables list. Reporting cadence, and what the report must contain, tells you what the agency believes it's on the hook for.
  • Ask who does the work. A senior pitch followed by junior execution is the standard failure mode. Ask for the names and the percentage of hours each will hold.
  • Ask about the transition out. If you left in month twelve, what stays with you? A real system stays. A set of ongoing tasks doesn't. This is the practical difference between a marketing vendor and an installed revenue system.

There's a related trap worth naming: hiring a strategist instead. The objection owners raise to us most often is "is this just a consultant who leaves when the contract's up?" — a fair question, and one we answer directly in why a fractional CMO doesn't fix what's actually broken. A plan handed over is not a system installed.


What should you check inside your own business before you evaluate anyone?

Audit what you already have. A large share of the revenue an agency will claim to "generate" is often sitting in your CRM as unworked leads and dead follow-up. If you buy new traffic before fixing that, you'll pay to widen a leaking pipe — and you won't be able to tell whose fault the leak is.

Before you take a single sales call:

  • Count your dormant leads. Everyone who enquired in the last 24 months and never bought. That list is usually the cheapest revenue in the building — see why reactivating old leads beats buying new ones.
  • Time your follow-up. How long between an enquiry and a first human or automated response? How many attempts before someone gives up? Run this follow-up audit first.
  • Establish your baseline numbers. Leads, contact rate, booking rate, close rate, average order value, repeat rate. Without a baseline, no agency can be held to anything. The revenue formula, broken down walks through each one.

Doing this changes the evaluation entirely. You stop asking agencies what they'd like to sell you and start asking which constraint they'd fix first — and why. The good ones diagnose. The rest recite their service list.


How do you evaluate an agency's use of AI without falling for the badge?

Ask what specific job the AI does and what number it moves. "AI-powered" as a label means nothing. AI doing reactivation outreach at 2 a.m., qualifying inbound leads before a human touches them, or drafting and testing follow-up at volume means something — because each of those has a measurable effect on booked revenue.

The follow-up question is the useful one: "Which of your AI components would you turn off tomorrow if it stopped moving the number?" An honest operator has an answer. Someone using AI as marketing decoration will change the subject.


Should you evaluate on price, and how do you compare offers fairly?

Compare total cost of the system, including your own time — not the invoice. Multi-vendor arrangements look cheaper line by line and cost more in aggregate. Vendor-management research puts coordination overhead at 8–15% of annual vendor spend in hidden internal labour that never appears on an invoice, with businesses reporting roughly 30% higher total spend versus a single integrated partner.

That hidden bill is real, and it lands on the owner: you become the integration layer between the SEO shop, the PPC shop, and the design shop, reconciling three versions of the message and three definitions of a lead. We've broken the arithmetic down in what using multiple marketing vendors actually costs.

Two fairness rules when comparing proposals:

  • Normalise the scope of outcome, not the scope of work. Two proposals with identical deliverables can carry completely different levels of accountability.
  • Price the ramp honestly. Real revenue systems take months to compound, not weeks. An agency that promises results in three weeks is either selling you a short-term paid-ads spike or something they can't install. Both end the same way.

What we've seen running these engagements

We install and run these systems for owner-led companies, so this checklist isn't theoretical — it's assembled from the questions clients wish they'd asked their previous agency, told to us in kickoff calls, month after month. Three patterns repeat with near-perfect consistency:

One. The client can name what the previous agency did but not what it built. Nothing was left behind because nothing was ever installed — the engagement was a subscription to activity.

Two. Nobody agreed on the number at the start. Twelve months in, both sides argue about whether it worked, because "worked" was never defined in dollars. Fixing this costs one conversation before signature and is nearly impossible afterwards.

Three. The CRM was the graveyard. Automations half-built, sequences that stopped after two emails, leads with no owner. Almost every engagement we start includes rebuilding this layer first — which is why we wrote what to install first in CRM automation, and what to skip.

If you take one thing from this page: the evaluation is not about choosing a supplier. It's about deciding whether a running machine will exist in your business when this is over, and who is on the hook for what it produces. Judge on that, and most of the shortlist eliminates itself.


The one-page diligence checklist

Print this. Take it into every call.

System

  • Named inventory of what is installed and running at day 90
  • What stays with me if the engagement ends
  • Which constraint they'd fix first, and their reasoning

Accountability

  • The single business number this engagement is held to
  • Named senior owner of that number, on the account, not just the pitch
  • Documented process for a month when the number falls

Proof

  • A real (redacted) client report from last month
  • Named client references I can actually call
  • Claims stated as facts they can evidence, not adjectives

Fit

  • Total cost including my coordination time
  • Honest ramp expectations, stated in months
  • Specific jobs AI performs and the number each one moves

Me

  • Dormant lead count established
  • Follow-up speed and attempt count measured
  • Baseline conversion and revenue metrics documented

Related reading: why a marketing plan isn't the same as a marketing system · net revenue: the number owner-operators should actually run on · what is deferred revenue, and what it tells an owner about the health of the machine


About the author

Avi Vatsa is CEO of Exchange Four Agency, where he leads the team that installs and runs AI-leveraged revenue systems for owner-led companies. His background spans law, technology, and marketing; he also co-founded Dialora, an AI voice-agent platform for automated lead capture and booking. Background sourced from public interviews on Marketer of the Day #1411 and the Jeremy Ryan Slate Show. Connect on LinkedIn.

Last reviewed: 22 August 2026.