For most Salesforce service teams, the safest first AI uses are case summaries and drafted replies that a person reviews before anything reaches a customer. Classification and routing come next, once you can check the model's suggestions against real case history. Customer-facing self-service agents belong last, limited to a few well-documented request types and a clean handoff to a person. Each step depends on current knowledge articles, consistent case data and someone reviewing the output on a schedule.
The five places AI fits in service
| Capability | What it does | What it depends on | Who sees the output first |
|---|---|---|---|
| Case and conversation summaries | Writes a summary, issue and resolution from the case or chat for the rep to edit and save | Complete case records and conversation transcripts | The rep |
| Reply drafting | Drafts email or chat replies from the conversation, optionally grounded in knowledge | Accurate articles and a defined tone of voice | The rep, who edits and sends |
| Knowledge grounding | Retrieves relevant articles or files so answers cite your content rather than general model knowledge | Articles that are current, owned and cover real questions | Whichever feature uses it |
| Classification and routing | Suggests or fills case fields based on similar closed cases, then routes on the result | Enough closed cases with reliable field values | The rep or supervisor, until trusted |
| Self-service agents | Answers customers and completes permitted actions on web, messaging or other channels | All of the above, plus actions and a handoff path | The customer |
Salesforce's own feature names here include Einstein Work Summaries, Einstein Service Replies, Einstein Case Classification and Agentforce for Service. Salesforce regroups and renames these features often, so check your contract for what is actually licensed before planning the order of rollout.
What to automate first
Rank candidate uses by two things: how much rep time they return, and how much harm a wrong output can do before someone catches it. Summaries score well on both, because a rep reads every one before saving it and a weak summary costs a minute of editing. Drafted replies save more time, and a person still controls what is sent. Classification errors are cheap if a rep confirms the value and expensive if a misrouted case sits in the wrong queue unnoticed.
Then look at where the workload concentrates. At a biotech company we worked with, one person on a three-person support team was handling 90% of the cases. Drafting email replies in the brand's tone, with recommended knowledge articles attached, went straight at that bottleneck, because the busiest person no longer started every answer from a blank screen. If your volume is spread evenly, a feature that helps every rep a little may matter more than one aimed at a single case type.
- Start with summaries on one queue, and compare edit rates across reps after a few weeks.
- Add reply drafting on the most common question type, internal-only, with reps editing every draft.
- Switch classification on in suggest mode, where reps accept or change the value, before letting it fill fields alone.
- Launch a self-service agent only for request types where drafts have been accepted with few changes.
- Hold back anything involving refunds, contract terms, safety or regulated advice until a person reviews each case.
Grounding answers in knowledge
Reply drafting and self-service agents are only as reliable as the content they retrieve. Service Replies can be grounded through Service AI Grounding or an Agentforce Data Library, drawing on Knowledge articles, uploaded files or the case itself. When a grounded draft is wrong, the cause is usually an outdated article, two articles that disagree, or a question no article answers.
Treat every rejected or heavily edited draft as a signal about the knowledge base. Add a simple way for reps to flag the article behind a bad answer, route those flags to the article owner, and track how many questions end with no relevant article found. That list is your content backlog, and it tells you which request types are ready for automation and which are not. Our Knowledge setup guide covers structure, audiences and upkeep.
Triage and routing
Einstein Case Classification recommends or fills case field values by learning from similar closed cases, and it can feed routing through assignment rules or Omni-Channel. The model learns from whatever your reps selected in the past, so a reason field that half the team left at its default will produce confident suggestions that are wrong. Audit a sample of closed cases for field accuracy before you train on them.
Keep routing logic readable even when AI fills the inputs. The model decides the case type or priority; your queues, skills and capacity settings decide who receives the work. That separation lets you switch the model off without rebuilding routing, and it lets supervisors see why a case landed where it did.
Self-service agents and the handoff
A customer-facing agent should know exactly which requests it handles and hand everything else to a person, with the context attached. Salesforce describes Agentforce for Service as able to escalate complex cases to service reps, and the quality of that moment decides how customers judge the whole experience. Pass the transcript, the case, what the agent already checked and why it escalated, so the customer never has to repeat themselves.
- Escalate on request: when a customer asks for a person, the agent hands over without arguing.
- Escalate on scope: anything outside the permitted request types goes to a queue, not a guess.
- Escalate on sentiment or risk: complaints, legal language or safety concerns go straight to staff.
- Escalate on failure: two unsuccessful attempts at the same question should trigger a handoff.
- Make handoffs visible in reporting so you can see which topics the agent should not own yet.
Quality control after launch
Assign an owner for AI output in service, usually a senior rep or team lead working with the admin, and give them a weekly review. The review should cover a random sample of summaries, sent replies and agent conversations, scored against a short rubric: accurate, complete, on-brand and within policy. Keep the scores in Salesforce so trends are reportable rather than anecdotal.
Compare cases handled with and without AI on the measures your customers feel: time to first response, reopen rate and satisfaction scores where you collect them. Expand to a new case type or channel only when the current one holds steady for several review cycles. When quality slips, look first at knowledge and case data, then at prompts and instructions, and only then at the model itself.
