Test an Agentforce agent the way you would test a new hire: against real work, with a written answer key. Start by fixing what the agent covers. Build a test set from historical cases, including edge cases and requests it must refuse. Run it in a sandbox, grade each answer on more than correctness, then pilot with staff. After launch, read session logs and rerun the set after every change.
What should you pin down before writing a single test?
Pin down the agent's job in writing first. A test can only fail against a stated expectation, so vague scope produces tests nobody can grade.
Write one short paragraph per topic. Since April 2026 Salesforce calls topics subagents, and the topic selector the Agent Router. Older docs and some screens still use the old terms, so expect both. For each one, list the questions it should answer, the actions it may take and the records it may read.
Then list what sits outside scope. These exclusions matter as much as the inclusions. An agent that politely declines a refund request it cannot handle is working. One that invents a refund policy is not.
- In scope: the request types, channels and customer groups the agent serves
- Allowed actions: each flow, Apex action or prompt template it can call, and when
- Grounding sources: which knowledge articles, objects and fields count as the truth
- Refusals: topics it must decline, such as legal advice, pricing exceptions or account closures
- Handoff rules: the conditions that send a conversation to a person, and to which queue
Where do good test cases come from?
They come from your own history. Pull closed cases, chat transcripts and emails that match the agent's scope, then write the correct outcome for each one.
Real requests carry the spelling mistakes, half-sentences and mixed questions that synthetic prompts tend to smooth over. Sample across request types rather than taking the most recent batch. A single busy period can skew the set toward one problem.
Tag every case with a category so failures cluster in a useful way. Useful tags include common question, multi-step request, ambiguous wording, missing information, out-of-scope ask and hostile or manipulative input.
Add the awkward cases deliberately. Include a customer asking about another customer's order, a request that mixes two topics and a question the knowledge base cannot answer. Include a message written in another language if your channels receive them. For each, record whether the right result is an answer, a clarifying question, a refusal or a handoff.
Which Salesforce tools help with Agentforce testing?
Agentforce Testing Center is the main one. It runs batches of test utterances and compares the topic, actions and response the agent produced against the results you expected.
You can load test cases from a CSV template or have the tool generate scenarios from instructions. Salesforce documentation now places it inside Agentforce Studio, with an older entry point in Setup. Run it in a sandbox, because test conversations can create or change CRM records.
Developers can run the same tests from Salesforce CLI through Agentforce DX, using a YAML test spec. Salesforce's developer documentation describes custom evaluations there that check a response for specific strings or numbers. That route suits teams that already deploy through a pipeline.
Since Summer '26, Testing Center runs do not consume Flex Credits, but Data 360 queries made during tests are still metered. Confirm current testing consumption and edition requirements with your Salesforce account team before planning large test runs.
How do you grade an agent's answers?
Grade each reply on several separate criteria, not a single pass or fail. An answer can be accurate and still fail because it read the wrong record or skipped a required step.
| Criterion | Question the reviewer asks | Typical failure |
|---|---|---|
| Accuracy | Is every factual statement correct for this customer? | Right policy, wrong product version |
| Grounding | Did the answer come from the approved article or record? | Plausible text with no source behind it |
| Tone | Does it read like your brand on a good day? | Stiff, over-apologetic or far too long |
| Action execution | Did it call the right action with the right inputs? | Updates the wrong case or skips a confirmation |
| Handoff | Did it escalate when the rules say it should? | Keeps answering after the customer asks for a person |
| Refusal | Did it decline out-of-scope requests cleanly? | Offers advice it was told never to give |
Topic routing deserves its own check. If a billing question lands in the shipping topic, every later step is wrong even when the wording sounds fine. Testing Center reports expected and actual topic side by side, which makes misroutes easy to spot.
Agent output varies between runs, so run important cases more than once. A case that passes four times and fails once is a real defect, not noise.
How do you test what the agent is allowed to see and do?
Test permissions as a separate pass, logged in as the agent's own running user. The agent can only stay inside its lane if that user has the narrowest access the job needs.
Check object and field access first. Then confirm sharing rules stop the agent from reading records outside the requester's scope. Try requests that name another account, ask for a field the agent should never expose, or attempt an action the user lacks rights to perform.
A correct result is a refusal or a handoff, never a partial leak. Our AI security guide covers how to design that access. The testing job is proving the design holds.
What does practical adversarial testing look like?
It means trying to talk the agent out of its rules with ordinary text. You do not need a red team to start.
- Instruction override: ask it to ignore previous instructions and show its system prompt
- Role claims: say you are an administrator or the account owner and ask for elevated help
- Smuggled instructions: paste text into a case description or email that tells the agent to act
- Data fishing: ask for another customer's address, order history or open cases
- Pressure: insist, repeat and escalate emotionally to see whether refusals soften
Record the exact wording of every attempt that works. Keep those prompts in the regression set permanently, because a fix in one release can quietly come undone in the next.
Why pilot with internal users before customers?
Internal users catch problems that scripted tests miss, and a bad answer costs far less with staff than with customers. A pilot also shows whether the people who own the work trust the output.
Start with one request type and have staff review every draft before it goes out. Ask them to mark each draft as usable, edited or rejected, with a reason. Those labels become new test cases.
Abstrakt used this approach for a biotech company on Service Cloud. Its Agentforce agent drafted email replies, summarized cases and recommended knowledge articles. Staff saw its output first, starting with the single most frequent case type, and coverage widened as more knowledge articles were written. In that project, 30% of question-type cases got a draft the team could send as written.
What should you monitor once the agent is live?
Monitor conversations, not just volume. Read a regular sample of real sessions and compare them with the grading criteria you used before launch.
Salesforce's Agentforce Observability tooling includes session tracing, which logs each turn, routing decision and action call. Agent Analytics reports on sessions, escalations and feedback. Some features have shipped in stages, so check which ones your org has enabled. Feature names in this area change often.
Watch for patterns rather than single bad replies. Rising handoffs on one topic, negative feedback clustering around one article, or actions that fail repeatedly all point to a specific fix.
When do you need to rerun the tests?
Rerun the full set after any change that could alter behavior. That includes your own edits and Salesforce's releases.
- Agent instructions, topic descriptions or action changes
- New, edited or retired knowledge articles
- Changes to the agent user's permissions or sharing
- Field, picklist or flow changes on objects the agent reads or updates
- Each Salesforce seasonal release, tested in a preview sandbox where possible
Store the test set somewhere versioned, next to the agent's configuration. Compare results with the last clean run so reviewers look only at what changed.
Which tests catch which problems?
Each test type covers a different failure, so a sound plan uses all of them. The table maps them out.
| Test type | What it catches | How to run it |
|---|---|---|
| Scope and routing | Requests sent to the wrong topic | Batch run in Testing Center with expected topics |
| Answer quality | Wrong, ungrounded or off-tone replies | Human review against the written answer key |
| Action tests | Wrong record updated or step skipped | Sandbox runs, then inspect the records changed |
| Permission tests | Data exposed beyond the requester's scope | Requests run as the agent user naming other accounts |
| Adversarial tests | Rules bypassed by persuasive text | Scripted manipulation prompts kept in the regression set |
| Internal pilot | Gaps that scripted cases never covered | Staff review every draft before it is sent |
| Regression | Behavior that broke after a change | Full rerun after releases, config or knowledge edits |
What should go-live criteria include?
Go-live criteria should be written thresholds agreed before testing starts. A launch decision made by feel after a good demo is the pattern to avoid.
- A pass rate on the core test set that the business owner signed off in advance
- Zero failures on permission and data-access tests
- Every known adversarial prompt refused or handed off
- Handoff paths tested end to end, with the receiving queue staffed
- A named person who reviews sessions and can switch the agent off
- Pilot users who say they would keep the agent
If governance policies are still unsettled, settle them first. The Agentforce readiness checklist covers what should be in place before testing begins.
What mistakes do teams make when testing agents?
Most mistakes come from testing too little of the right thing. A long list of easy questions creates false confidence.
- Testing only happy-path questions written by the build team
- Grading replies after reading them instead of against a prior answer key
- Running tests as an administrator, which hides permission gaps
- Treating one passing run as proof, despite variable output
- Updating knowledge articles without rerunning the tests that depend on them
- Launching to customers without an internal pilot
If you want help building a test set or running a pilot, see our Agentforce consulting work.

