Calculator and pen on a page of handwritten figures

Photo: Aaron Lefler / Unsplash

Guide

How to build an honest business case for Agentforce

A method for CFOs and COOs to build an Agentforce business case: pick a countable use case, baseline your own costs, model ranges, load every cost, price risk and run a pilot with go/no-go criteria.

An honest Agentforce business case starts with one use case you can count, such as support cases or lead follow-ups. Baseline what each unit costs today from your own records. Model agent resolution as a range you intend to test, not a promise. Load every cost, including usage, data and knowledge work, build, testing and ongoing tuning. Then let a scored pilot decide whether to scale.

Why do most AI agent business cases fall apart in review?

They usually borrow a vendor benchmark instead of measuring the company's own work. A finance reviewer then asks where the deflection figure came from, and nobody can trace it.

The second failure is a cost side that stops at the subscription. Data preparation, knowledge cleanup, testing and the person who reviews agent conversations every week rarely appear. The third is treating a hoped-for outcome as a line item. A credible case reads more like a test plan with a budget attached.

This guide covers the method a CFO or COO can defend. For the Salesforce-wide ROI formula, see our guide on calculating Salesforce ROI. For a breakdown of what drives Agentforce spend, see the cost guide linked at the end.

Which use case should the business case cover first?

Pick one workflow with high, steady volume, a clear definition of done and records already in Salesforce. Breadth can come later; the first case has to be provable.

Good candidates share a few traits. Each unit of work is logged as a record, so you can count it. The work repeats in recognizable patterns, and the answers already exist somewhere, such as knowledge articles or order data. A wrong answer is recoverable, because a person can catch it before harm is done.

  • Service: inbound question-type cases, order status inquiries, password or access requests, returns eligibility.
  • Sales: first follow-up on inbound leads, meeting scheduling, routing unqualified inquiries.
  • Internal: employee questions about policies, benefits or IT that land in a shared queue.

Avoid starting with work that is rare, judgment-heavy, regulated in ways nobody has mapped, or tracked mainly in email and spreadsheets. Those can work later, but they make poor evidence.

How do you baseline the current cost of each unit of work?

Pull volume and handling time from Salesforce reports, then multiply time by a loaded labor rate your finance team already uses. Do this before any build starts, because the baseline cannot be reconstructed afterward.

For a service use case, gather several months of case data by type. Capture volume, average handle time, first-contact resolution, reopen rate and escalations. If handle time is not tracked, sample it: have agents time a set of cases for two representative periods. Imperfect but documented beats precise but invented.

Translate time into money with the fully loaded cost per hour finance uses for headcount planning. Add any per-contact costs, such as outsourced overflow or telephony. Keep the arithmetic visible so a reviewer can change one input and see the effect.

Also baseline quality. Customer satisfaction, error rates and compliance findings matter because an agent that cuts cost while raising complaints is not a saving.

What counts as "resolved by the agent"?

Write the definition down before the pilot starts and get service and finance leaders to sign it. Without it, every result becomes an argument.

Separate outcomes into distinct categories, because they carry different value:

  • Fully resolved: the customer or employee got a correct answer and did not come back on the same issue within an agreed window.
  • Assisted: the agent drafted a reply, summary or next step that a person reviewed and sent, with little or no editing.
  • Escalated cleanly: the agent handed off with useful context, so the person started ahead of where they would have.
  • Failed: wrong answer, unhelpful loop, or an escalation that lost context and made the case longer.

Then decide how you will audit quality. A common pattern is a weekly sample of agent conversations scored against a short rubric: accuracy, tone, policy compliance and whether escalation happened when it should. Name the reviewer and the sample size in the plan.

How should deflection and assist rates appear in the model?

As ranges with a low, expected and high case, each labeled as an assumption the pilot will test. Never as a single point taken from a vendor slide.

Build the model so savings equal volume times the share handled, times time saved per unit, times loaded rate. Run it at all three levels. If the business case only works at the high end, that is useful information before you spend anything.

Treat assist differently from full resolution. A drafted reply that saves part of a person's handling time is real value, but it rarely removes a role. Most early gains show up as capacity: the same team absorbing more volume, or spending time on harder work.

Which costs belong in an Agentforce business case?

All of them: platform usage, the data and knowledge work underneath, the build, and the people who run the agent after launch. The subscription is often the smaller and more predictable part.

Salesforce currently offers several pricing structures for Agentforce. These include consumption through Flex Credits, a per-conversation model, and per-user add-ons or editions. Packaging changes often, so confirm the current options on salesforce.com and with your Salesforce account team. Model usage against your own expected volume rather than a list example.

  • Usage: Flex Credits or conversations consumed at pilot volume and at full volume, plus any add-on or edition fees.
  • Data preparation: deduplication, missing fields and record ownership the agent will rely on.
  • Data Cloud (Data 360), if the use case needs data from outside the core CRM. Confirm with your account team whether it is required and how it is licensed for your scenario.
  • Knowledge cleanup: retiring stale articles, filling gaps and assigning owners to keep them current.
  • Build: subagents (formerly topics), instructions, actions, Flows or Apex the agent calls, and any integrations.
  • Testing: test scenarios, sandbox time and the people who review results before launch.
  • Run costs: conversation review, tuning, knowledge upkeep, monitoring usage and governance meetings.

How do you separate hard savings from soft benefits?

Hard savings change a budget line you can point to; soft benefits improve something real but do not reduce spend on their own. Report both, but total them separately.

Hard savings include avoided hires against a documented growth plan, reduced outsourced overflow and lower overtime. Soft benefits include faster response times, longer service hours from self-service, more consistent answers and staff time moved to complex work. A CFO will usually fund on hard savings and treat soft benefits as upside.

Be careful with "hours freed." Freed hours are only a saving if you plan what happens to them. Write down whether they absorb growth, shift to higher-value work, or reduce future hiring.

How do you price the risks of a wrong answer?

Estimate the cost of each failure type, multiply by a plausible failure rate, and subtract it from the benefit. Then design controls that lower the rate and include their cost too.

The main risks are predictable. A wrong answer can cause a refund, a complaint or a compliance issue. A poor escalation can make a case take longer than if a person had handled it from the start. Customers who dislike the experience may contact you again through another channel, which inflates volume.

Price these with your own data where possible. Use the average cost of a complaint, a refund or a repeat contact. Controls include restricting the agent to approved knowledge, requiring human review for certain topics and keeping clear escalation paths. Salesforce's testing and monitoring tools and the Einstein Trust Layer help, but they do not replace your own review.

What should an Agentforce pilot look like?

Narrow scope, a fixed measurement window, a control group where possible, and go/no-go criteria agreed before launch. The pilot exists to replace your assumptions with measured ranges.

Start internally if risk is a concern: the agent drafts, people send. That produces measurable data on draft quality without exposing customers. Expand to customer-facing use only when the internal results clear the bar.

Set criteria in advance. Examples include a minimum share of usable drafts or resolutions and a maximum failure rate in quality review. Add no rise in reopen or complaint rates, and usage cost within the modeled range. If the pilot misses, decide whether to fix the inputs, such as knowledge or data, or stop.

What does a usable business case worksheet look like?

Each line item needs a measurement method and an owner. If no one owns a number, it will not be measured, and the case cannot be checked later.

Agentforce business case worksheet: line items, measurement and ownership
Line itemHow to measure itOwner
Monthly volume for the use caseSalesforce report by case type, lead source or queue over recent monthsService or sales operations
Handle time per unitTracked time fields, or a timed sample where none existTeam lead for the queue
Loaded labor rateFinance's standard fully loaded cost per hourFinance
Resolution and assist ratesLow, expected and high range, replaced by pilot resultsPilot owner
Quality scoreWeekly scored sample of agent conversations against a written rubricQuality or service lead
Usage costConsumption or conversations at pilot and projected volume, using current contract termsSalesforce admin with finance
Data and knowledge workHours estimated for cleanup and gap filling, then trackedData owner and knowledge manager
Build and testingPartner or internal estimate for actions, Flows and test cyclesProject sponsor
Ongoing tuning and governanceRecurring hours for review, updates and oversight meetingsAgent product owner
Risk allowanceExpected failures multiplied by average cost per complaint, refund or repeat contactFinance with service lead
Hard savingsAvoided hires, overflow or overtime tied to a dated planCOO or department head
Soft benefitsResponse time, coverage and satisfaction, reported separatelyService or sales leader

When should you expand beyond the first use case?

When the pilot meets its criteria for a sustained period and the run costs are known, not estimated. Then repeat the same method for the next workflow.

Each new use case gets its own baseline, definition of resolution and go/no-go criteria. Reusing the first case's numbers for a different workflow is how optimistic estimates creep back in. Our Agentforce consultants can help scope a first use case and design the pilot measures. We do this work across service, sales and internal support.

Chris Gooding, President & CEO of Abstrakt Solutions
President & CEO, Abstrakt Solutions
LinkedIn →

Tech Talk

A monthly brief for the people who own Salesforce, AI and revenue technology

What changed in Salesforce and AI this month, and what to do about it.

One email a month. Written by the consultants who deliver the work, not by a marketing team, for the leaders who make the technology decisions.

  • What changed in Salesforce, AI, integration and RevOps, and what it means for your org
  • At least one framework, checklist or reference architecture you can take into a meeting
  • Honest opinions, including when we disagree with what a vendor is selling
  • No sales sequence. We do not sell from this list

Consultant analysis, not vendor recaps. One click to leave.

One email a month. Your industry and your address, nothing else. We never share either, and you can unsubscribe from the bottom of any issue. See what’s in Tech Talk →

Call (314) 916-4095 Book a consultation
Call (314) 916-4095 Book a call