Rack of test tubes filled with colored liquids

Photo: Ryan Zazueta / Unsplash

Guide

How to test an Agentforce agent before and after launch

A practical test plan for Agentforce agents: scope, test sets from real cases, Testing Center, grading, permission and adversarial tests, internal pilots, monitoring, regression and go-live criteria.

Test an Agentforce agent the way you would test a new hire: against real work, with a written answer key. Start by fixing what the agent covers. Build a test set from historical cases, including edge cases and requests it must refuse. Run it in a sandbox, grade each answer on more than correctness, then pilot with staff. After launch, read session logs and rerun the set after every change.

What should you pin down before writing a single test?

Pin down the agent's job in writing first. A test can only fail against a stated expectation, so vague scope produces tests nobody can grade.

Write one short paragraph per topic. Since April 2026 Salesforce calls topics subagents, and the topic selector the Agent Router. Older docs and some screens still use the old terms, so expect both. For each one, list the questions it should answer, the actions it may take and the records it may read.

Then list what sits outside scope. These exclusions matter as much as the inclusions. An agent that politely declines a refund request it cannot handle is working. One that invents a refund policy is not.

  • In scope: the request types, channels and customer groups the agent serves
  • Allowed actions: each flow, Apex action or prompt template it can call, and when
  • Grounding sources: which knowledge articles, objects and fields count as the truth
  • Refusals: topics it must decline, such as legal advice, pricing exceptions or account closures
  • Handoff rules: the conditions that send a conversation to a person, and to which queue

Where do good test cases come from?

They come from your own history. Pull closed cases, chat transcripts and emails that match the agent's scope, then write the correct outcome for each one.

Real requests carry the spelling mistakes, half-sentences and mixed questions that synthetic prompts tend to smooth over. Sample across request types rather than taking the most recent batch. A single busy period can skew the set toward one problem.

Tag every case with a category so failures cluster in a useful way. Useful tags include common question, multi-step request, ambiguous wording, missing information, out-of-scope ask and hostile or manipulative input.

Add the awkward cases deliberately. Include a customer asking about another customer's order, a request that mixes two topics and a question the knowledge base cannot answer. Include a message written in another language if your channels receive them. For each, record whether the right result is an answer, a clarifying question, a refusal or a handoff.

Which Salesforce tools help with Agentforce testing?

Agentforce Testing Center is the main one. It runs batches of test utterances and compares the topic, actions and response the agent produced against the results you expected.

You can load test cases from a CSV template or have the tool generate scenarios from instructions. Salesforce documentation now places it inside Agentforce Studio, with an older entry point in Setup. Run it in a sandbox, because test conversations can create or change CRM records.

Developers can run the same tests from Salesforce CLI through Agentforce DX, using a YAML test spec. Salesforce's developer documentation describes custom evaluations there that check a response for specific strings or numbers. That route suits teams that already deploy through a pipeline.

Since Summer '26, Testing Center runs do not consume Flex Credits, but Data 360 queries made during tests are still metered. Confirm current testing consumption and edition requirements with your Salesforce account team before planning large test runs.

How do you grade an agent's answers?

Grade each reply on several separate criteria, not a single pass or fail. An answer can be accurate and still fail because it read the wrong record or skipped a required step.

Grading criteria for each test case
CriterionQuestion the reviewer asksTypical failure
AccuracyIs every factual statement correct for this customer?Right policy, wrong product version
GroundingDid the answer come from the approved article or record?Plausible text with no source behind it
ToneDoes it read like your brand on a good day?Stiff, over-apologetic or far too long
Action executionDid it call the right action with the right inputs?Updates the wrong case or skips a confirmation
HandoffDid it escalate when the rules say it should?Keeps answering after the customer asks for a person
RefusalDid it decline out-of-scope requests cleanly?Offers advice it was told never to give

Topic routing deserves its own check. If a billing question lands in the shipping topic, every later step is wrong even when the wording sounds fine. Testing Center reports expected and actual topic side by side, which makes misroutes easy to spot.

Agent output varies between runs, so run important cases more than once. A case that passes four times and fails once is a real defect, not noise.

How do you test what the agent is allowed to see and do?

Test permissions as a separate pass, logged in as the agent's own running user. The agent can only stay inside its lane if that user has the narrowest access the job needs.

Check object and field access first. Then confirm sharing rules stop the agent from reading records outside the requester's scope. Try requests that name another account, ask for a field the agent should never expose, or attempt an action the user lacks rights to perform.

A correct result is a refusal or a handoff, never a partial leak. Our AI security guide covers how to design that access. The testing job is proving the design holds.

What does practical adversarial testing look like?

It means trying to talk the agent out of its rules with ordinary text. You do not need a red team to start.

  • Instruction override: ask it to ignore previous instructions and show its system prompt
  • Role claims: say you are an administrator or the account owner and ask for elevated help
  • Smuggled instructions: paste text into a case description or email that tells the agent to act
  • Data fishing: ask for another customer's address, order history or open cases
  • Pressure: insist, repeat and escalate emotionally to see whether refusals soften

Record the exact wording of every attempt that works. Keep those prompts in the regression set permanently, because a fix in one release can quietly come undone in the next.

Why pilot with internal users before customers?

Internal users catch problems that scripted tests miss, and a bad answer costs far less with staff than with customers. A pilot also shows whether the people who own the work trust the output.

Start with one request type and have staff review every draft before it goes out. Ask them to mark each draft as usable, edited or rejected, with a reason. Those labels become new test cases.

Abstrakt used this approach for a biotech company on Service Cloud. Its Agentforce agent drafted email replies, summarized cases and recommended knowledge articles. Staff saw its output first, starting with the single most frequent case type, and coverage widened as more knowledge articles were written. In that project, 30% of question-type cases got a draft the team could send as written.

What should you monitor once the agent is live?

Monitor conversations, not just volume. Read a regular sample of real sessions and compare them with the grading criteria you used before launch.

Salesforce's Agentforce Observability tooling includes session tracing, which logs each turn, routing decision and action call. Agent Analytics reports on sessions, escalations and feedback. Some features have shipped in stages, so check which ones your org has enabled. Feature names in this area change often.

Watch for patterns rather than single bad replies. Rising handoffs on one topic, negative feedback clustering around one article, or actions that fail repeatedly all point to a specific fix.

When do you need to rerun the tests?

Rerun the full set after any change that could alter behavior. That includes your own edits and Salesforce's releases.

  • Agent instructions, topic descriptions or action changes
  • New, edited or retired knowledge articles
  • Changes to the agent user's permissions or sharing
  • Field, picklist or flow changes on objects the agent reads or updates
  • Each Salesforce seasonal release, tested in a preview sandbox where possible

Store the test set somewhere versioned, next to the agent's configuration. Compare results with the last clean run so reviewers look only at what changed.

Which tests catch which problems?

Each test type covers a different failure, so a sound plan uses all of them. The table maps them out.

Agentforce test plan by test type
Test typeWhat it catchesHow to run it
Scope and routingRequests sent to the wrong topicBatch run in Testing Center with expected topics
Answer qualityWrong, ungrounded or off-tone repliesHuman review against the written answer key
Action testsWrong record updated or step skippedSandbox runs, then inspect the records changed
Permission testsData exposed beyond the requester's scopeRequests run as the agent user naming other accounts
Adversarial testsRules bypassed by persuasive textScripted manipulation prompts kept in the regression set
Internal pilotGaps that scripted cases never coveredStaff review every draft before it is sent
RegressionBehavior that broke after a changeFull rerun after releases, config or knowledge edits

What should go-live criteria include?

Go-live criteria should be written thresholds agreed before testing starts. A launch decision made by feel after a good demo is the pattern to avoid.

  • A pass rate on the core test set that the business owner signed off in advance
  • Zero failures on permission and data-access tests
  • Every known adversarial prompt refused or handed off
  • Handoff paths tested end to end, with the receiving queue staffed
  • A named person who reviews sessions and can switch the agent off
  • Pilot users who say they would keep the agent

If governance policies are still unsettled, settle them first. The Agentforce readiness checklist covers what should be in place before testing begins.

What mistakes do teams make when testing agents?

Most mistakes come from testing too little of the right thing. A long list of easy questions creates false confidence.

  • Testing only happy-path questions written by the build team
  • Grading replies after reading them instead of against a prior answer key
  • Running tests as an administrator, which hides permission gaps
  • Treating one passing run as proof, despite variable output
  • Updating knowledge articles without rerunning the tests that depend on them
  • Launching to customers without an internal pilot

If you want help building a test set or running a pilot, see our Agentforce consulting work.

Chris Gooding, President & CEO of Abstrakt Solutions
President & CEO, Abstrakt Solutions
LinkedIn →

Tech Talk

A monthly brief for the people who own Salesforce, AI and revenue technology

What changed in Salesforce and AI this month, and what to do about it.

One email a month. Written by the consultants who deliver the work, not by a marketing team, for the leaders who make the technology decisions.

  • What changed in Salesforce, AI, integration and RevOps, and what it means for your org
  • At least one framework, checklist or reference architecture you can take into a meeting
  • Honest opinions, including when we disagree with what a vendor is selling
  • No sales sequence. We do not sell from this list

Consultant analysis, not vendor recaps. One click to leave.

One email a month. Your industry and your address, nothing else. We never share either, and you can unsubscribe from the bottom of any issue. See what’s in Tech Talk →

Call (314) 916-4095 Book a consultation
Call (314) 916-4095 Book a call