Checked against Salesforce Help, the Salesforce Developers documentation and Trailhead, September 2026. Data Cloud is now called Data 360.
Most architecture slides draw the Einstein Trust Layer as one box between Salesforce and the model — OpenAI, Claude or Gemini. That box hides the question a compliance lead in a pharma company actually asks: what exactly leaves our org, and who sees it? The answer depends on which of seven steps run for a given call — and, since 2025, on whether the call comes from a prompt template or
The scenario
Norvane Pharma (fictional) runs its medical information desk on Service Cloud, with Data 360 underneath. Two generative AI use cases are on the table: a case summary button for the inquiry team, and later an Agentforce service agent on the HCP portal. Both call a large language model. Both ground on the same unified HCP profile. They do not get the same protection — and the difference is documented, just not on the slide.
The Einstein Trust Layer round trip, step by step
- Secure data retrieval and grounding. Prompt Builder resolves merge fields from record fields, related lists, Flow or Apex outputs, Data 360 DMOs and retrievers. Retrieval runs with the permissions of the user executing the prompt; role-based controls and field-level security stay in place.
- Data masking. Pattern-based detection (machine learning for names of people and companies, regex patterns plus context words for email, phone, credit card, IBAN and passport numbers) and field-based masking for fields tagged with Shield Platform Encryption or data classification. Detected values become placeholders; the Trust Layer temporarily stores the mapping.
- Prompt defense. System policies are added to the prompt to resist jailbreaking and prompt injection, and to keep the model from answering where it has no information. Separately, Prompt Injection Detection (Beta) scores prompts for injection attempts and reports them in Data 360 — currently for English (United States) only.
- LLM gateway and model. The request goes to the selected model. For partner models such as OpenAI, zero data retention applies: data is deleted after the response is sent back and is not used for training.
- Toxicity detection. The response is scored per category — hate, violence, physical harm, sexual content, profanity and further categories that differ slightly between APIs — plus an overall safety score. Input safety scoring can be switched on for the prompt as well. The documentation describes scoring and storing — not blocking.
- Demasking. Placeholders are swapped back using the mapping saved in step 2.
- Feedback and audit. The original and the masked prompt, the raw and the demasked response, toxicity scores and user feedback land in Data 360 as a timestamped audit trail.
Step 1 in detail: grounding decides what can leave
Grounding is dynamic: the template is resolved at run time, for the user who triggers it. Trailhead calls this dynamic grounding, and it is why the same template produces different prompts for different users. What reaches the model depends on which data provider fills each part of the template:
- Record merge fields and related lists — data from the input record and its relationships. Related lists follow the page layout of the running user, record-level filters are not applied, and the result is rendered as JSON. If the user lacks the related list on the layout or Read access to the object, no related list data is sent to the model at all.
- Flow — a Template-Triggered Prompt Flow receives the template inputs and adds text through the Add Prompt Instructions element. It can pull CRM data, Data 360 data and external sources.
- Apex — an invocable method bound to the template type returns a
Promptstring. Same reach as Flow, with full code control. - Retrievers — semantic search over indexed, unstructured content such as knowledge articles. The retrieved chunks are inserted into the prompt: this is the RAG part of the pipeline.
An Apex grounding class, following the pattern in the Salesforce Developers blog. The class decides exactly which fields become prompt text — a useful control point when a use case must not send certain data:
public with sharing class HcpInquiryContext {
// Illustrative: bound to a Flex template via its API name.
// Fields used here must be available on the input record.
@InvocableMethod(capabilityType='FlexTemplate://Case_Summary')
public static List<Response> build(List<Request> requests) {
Case c = requests[0].caseRecord;
// Only what is added here becomes prompt text
Response res = new Response();
res.Prompt = 'Product: ' + c.Product__c
+ '\nInquiry type: ' + c.Type
+ '\nOpened: ' + String.valueOf(c.CreatedDate.date());
return new List<Response>{ res };
}
public class Request {
@InvocableVariable public Case caseRecord;
}
public class Response {
@InvocableVariable public String Prompt;
}
}
Step 2 in detail: what Salesforce data masking actually does
Masking replaces each detected value with a placeholder that says what it was, so the model keeps the context of the sentence without seeing the value. For Norvane’s case summary, the difference looks like this — values invented, placeholder format as shown in Salesforce’s Trailhead example:
Before (in the org) After (sent to the model)
------------------------------------ ------------------------------------
Caller: Dr. Anna Weber Caller: <NAME_1>
Clinic: Praxis Weber Hamburg Clinic: <COMPANY_1>
Email: a.weber@example.de Email: <EMAIL_1>
Phone: +49 40 1234 5678 Phone: <PHONE_1>
Product: Norvastin 20 mg Product: Norvastin 20 mg
Three details matter for European orgs:
- Language coverage. Name, company name, email, phone, credit card, IBAN and passport numbers are detected in English, French, German, Italian, Japanese and Spanish. Driver’s licence, ITIN and Social Security numbers are US English only. Product names, diagnoses or free-text symptoms are not a masking category.
- Field-based masking has a narrow scope. It supports only record merge fields and related lists. Data that enters the prompt through Flow, Apex or a retriever gets pattern-based detection only — a classified field is not protected by its classification once Apex has turned it into prompt text.
- Locale. For direct Models API calls,
localization.defaultLocaledetermines the region used for data masking. If no localization is passed, Salesforce assumesen_US— so German phone numbers and other region-specific formats may be detected less reliably.
The detail that changes the design: no data masking for Agentforce agents
“In Einstein Trust Layer, pattern-based and field-based data masking for large language models (LLMs) is disabled for agents.”
Salesforce Help
Salesforce gives two reasons: accuracy — an agent asked to build a list of accounts similar to a reference account fails if it only sees that account as a placeholder — and performance. Both are reasonable. The consequence is still worth stating plainly.
For Norvane, the case summary button sends the model placeholders instead of the HCP’s name and contact details. The portal agent sends them in clear text. The same prompt template behaves differently depending on where it runs: called from a Lightning page or a Flow it follows the masking settings; called as an agent action it doesn’t. Masking remains available for embedded features such as Einstein Service Replies and Work Summaries.
The model behind the agent matters too. Agentforce’s Salesforce Default option is a managed mix of models that currently includes GPT-4o. The AWS-Hosted option uses an Anthropic Claude model on Amazon Bedrock. Those two options sit in different trust zones.
For Agentforce security, protection moves from the model never sees the value to two other levers: where the model runs and what the contract with its provider says. That turns model selection from a quality decision into a data-flow decision.
Why OpenAI is in the picture — and why it’s one of three zones
Salesforce builds its own models, but for open-ended language tasks the frontier models are stronger and improve faster than a CRM vendor could retrain. So Salesforce buys inference. The Trust Layer is the condition under which that is acceptable. What most people miss is that “external model” covers three different trust arrangements.
- Salesforce trust boundary. Models built by Salesforce, and — less obviously — Anthropic and Amazon models on Amazon Bedrock, which Salesforce describes as operated on Amazon Bedrock infrastructure “entirely within the Salesforce Trust Boundary”. For agents, Salesforce Help adds that data sent to such a model doesn’t leave the Salesforce trust boundary.
- Shared trust boundary. Models operated by Salesforce partners, such as OpenAI and Azure OpenAI, covered by Salesforce’s zero data retention policy: data isn’t retained and is deleted after the response is sent back.
- Bring your own LLM. A model you connect through AI Models with your own credentials, on Amazon Bedrock, Azure OpenAI, OpenAI or Vertex AI. Requests still route through the Models API, and Trust Layer features are supported. But Salesforce’s Agentforce privacy FAQ states that it does not apply to BYO LLM — your own contract with the provider governs retention.
For Norvane’s portal agent, which grounds on HCP data without masking, the question “which model?” therefore has a second half: “in which zone?”
Why zero data retention moves moderation into Salesforce
Trailhead explains a detail that is easy to miss. Model providers normally keep prompts and responses for a while to monitor abuse. Salesforce’s zero data retention agreement removes that: the provider forgets prompt and response once the answer is back. The moderation the provider would otherwise do has to happen somewhere, so Salesforce does it inside the Trust Layer — which is why toxicity scoring and the audit trail live on the Salesforce side, in your Data 360.
The way back: toxicity, citations, feedback
- Toxicity scoring runs on the response before demasking. The model returns an
isToxicityDetectedflag when it is highly confident;falseonly means nothing was detected. Salesforce also notes that free-text input can carry toxic language from end users, so prompts can be scored too. - Citations link a generated answer to the source records or articles it used, so users can verify it. They are supported in English, French, German, Italian, Brazilian Portuguese and Spanish.
- Feedback comes in two forms. Explicit: thumbs up or down, with reasons such as factually incorrect, incomplete, biased or harmful, or wrong tone. Implicit, depending on the feature: whether the user accepted, edited or discarded the draft. Both land in the audit data.
What it looks like in code
Running a prompt template from Apex, following the pattern in Salesforce’s Apex reference:
public with sharing class CaseSummaryService {
public static String summarize(Id caseId) {
// Record inputs are passed as a map with an 'id' key
Map<String, String> caseRef = new Map<String, String>{ 'id' => caseId };
ConnectApi.WrappedValue caseValue = new ConnectApi.WrappedValue();
caseValue.value = caseRef;
Map<String, ConnectApi.WrappedValue> params =
new Map<String, ConnectApi.WrappedValue>();
params.put('Input:Case', caseValue);
ConnectApi.EinsteinPromptTemplateGenerationsInput input =
new ConnectApi.EinsteinPromptTemplateGenerationsInput();
input.inputParams = params;
input.isPreview = false;
input.additionalConfig = new ConnectApi.EinsteinLlmAdditionalConfigInput();
input.additionalConfig.applicationName = 'PromptTemplateGenerationsInvocable';
// Everything from here is platform behaviour:
// grounding, masking, prompt defense, model call,
// toxicity scoring, demasking, audit
ConnectApi.EinsteinPromptTemplateGenerationsRepresentation result =
ConnectApi.EinsteinLLM.generateMessagesForPromptTemplate(
'Case_Summary', input);
return result.generations[0].text;
}
}
What the code doesn’t contain is the point: no named credential to OpenAI, no API key, no masking call, no logging. Six of the seven steps are platform behaviour. The response carries the toxicity result alongside the text — field names as documented in the Apex reference, values illustrative:
{
"promptTemplateDevName": "Case_Summary",
"generations": [
{
"text": "Dr. Weber asked on 12 September for ...",
"contentQualityRepresentation": { "isToxicityDetected": false },
"safetyScoreRepresentation": {
"safetyScore": 0.999,
"toxicityScore": 0.0004,
"hateScore": 0.0001,
"violenceScore": 0.0002,
"physicalScore": 0.0,
"sexualScore": 0.0,
"profanityScore": 0.0003
}
}
]
}
Category scores run from 0 to 1, with 1 the most toxic and 0.5 the threshold. The safety score runs the other way: 1 is safest, 0.5 and above counts as safe.
Without a template, the Models API calls a model directly — still through the gateway and the Trust Layer. The request below uses the documented endpoint and headers and passes a German locale so masking and toxicity scoring use the right region:
POST https://api.salesforce.com/einstein/platform/v1/models/sfdc_ai__DefaultBedrockAnthropicClaude45Sonnet/generations
Authorization: Bearer <JWT>
Content-Type: application/json
x-sfdc-app-context: EinsteinGPT
x-client-feature-id: ai-platform-models-connected-app
{
"prompt": "Summarise this inquiry in three sentences: ...",
"localization": {
"defaultLocale": "de_DE",
"expectedLocales": ["de_DE"]
}
}
Here the toxicity result comes back as generation.contentQuality.scanToxicity, with isDetected and a list of category scores. There is no programmatic way to read the masked version from the Models API; that lives in the audit trail.
Where the architecture work actually sits
The gateway and the scoring models are platform. The decisions that shape what reaches a model are yours:
- Grounding. Merge fields, Flow and Apex grounding, Data 360 DMOs and retrievers decide what goes into the prompt — and the running user’s access decides what can. In Data 360, identity resolution decides whether the model sees one HCP or five partial duplicates.
- Masking configuration. Which data types are on. At initial setup, the most commonly used types are on and less frequent ones are off. Every type you mask costs the model context — and with masking turned on, all models are limited to a 65,536-token context window. A trade-off to agree with the business, not a default to accept.
- Model selection by zone. Especially for agents, where masking doesn’t apply.
- Audit data. Audit and feedback records live in your Data 360 instance, where you control how long they are kept; Salesforce additionally stores them for 30 days for compliance purposes. They map to standard DMOs — Ai Gateway Request, Ai Response Generation, Ai Content Quality, Ai Feedback — and Salesforce recommends building reports on those rather than the legacy GenAI objects. Prompt Injection Detection (Beta) adds a monitoring view on the same data.
| Standard DMO | Legacy DMO | What it holds |
|---|---|---|
| Ai Gateway Request | GenAIGatewayRequest | Prompt before masking, masked prompt, model, provider, template name and version, masking and safety-scoring flags |
| Ai Response Generation | GenAIGeneration | Generated response |
| Ai Content Quality | GenAIContentQuality | Toxicity flag, per prompt (input) or response (output) |
| Ai Content Quality Category | GenAIContentCategory | Detector results by category: PII counts, toxicity scores, prompt defense |
| Ai Feedback | GenAIFeedback | Accept, edit, reject, thumbs up or down |
| Ai Feedback Additional Info | GenAIFeedbackDetail | Feedback reasons and free text |
What the Trust Layer does not do
- It does not mask data for Agentforce agents.
- It does not check facts. Grounding and system policies reduce invented answers; nothing verifies them.
- It does not replace permissions. If the running user can see a field, the prompt can contain it.
- It does not guarantee detection. Salesforce states that no model reaches 100% accuracy.
- It does not document blocking. Toxicity and prompt injection detection score and record.
- It does not keep sensitive data out of the audit trail. The Gateway Request object stores the prompt before masking, so access to the audit data needs its own governance.
Questions to settle before the first prompt goes live
- Which use cases run as agents, and which as embedded features or prompt templates? Only the second group is masked.
- For each use case, which trust zone does the selected model sit in?
- Which fields are grounded, and does the running user’s access match what should reach a model?
- Which masking types are on, and which are deliberately off because they break the output?
- Which data enters through Flow, Apex or retrievers and therefore gets pattern-based masking only?
- Do API integrations pass the right locale, or does masking run on the
en_USdefault? - If BYOLLM is on the table, who owns the provider contract and its retention terms?
- Who may see the audit and feedback data in Data 360, who reviews it, and how long is it kept?
None of this makes the Trust Layer weaker than the slides suggest. It makes it specific. The one-box version lets a steering committee approve “AI with guardrails”. The seven-step version tells them which guardrail applies to which use case — and where the remaining decisions sit with them.
FAQ: Salesforce and LLMs
Does Salesforce use OpenAI?
Yes, among others. OpenAI and Azure OpenAI models are Salesforce-managed models in the shared trust boundary, covered by Salesforce’s zero data retention policy. Salesforce also offers Anthropic and Amazon models on Amazon Bedrock, Google Gemini models on Vertex AI and its own models, and customers can connect their own model through AI Models.
Which LLM does Agentforce use?
By default, the Salesforce Default option: a managed mix of models that currently includes GPT-4o. The AWS-Hosted option uses an Anthropic Claude model on Amazon Bedrock, inside the Salesforce trust boundary. Custom actions that run prompt templates, Apex or the Models API can use any Salesforce-managed or bring-your-own model.
Does Salesforce send customer data to ChatGPT?
Not to the ChatGPT app. Prompts can go to OpenAI models through the Salesforce LLM gateway, and grounded prompts can contain customer data. Under Salesforce’s zero data retention policy, that data is not retained or used for training and is deleted after the response is sent back. For prompt templates, sensitive values can be masked first; for Agentforce agents they are not.
How does Agentforce protect data sent to an LLM?
Data masking is disabled for agents, so protection comes from other layers: grounding only with data the running user may see, the trust zone of the selected model, zero data retention, prompt defense, toxicity scoring and an audit trail in Data 360. Choosing a model inside the Salesforce trust boundary keeps the data there.
Also: Could this happen to us? When sensitive customer data reaches the wrong department — who may see which data inside Salesforce Data 360.
- Trailhead — The Agentforce Trust Layer: Follow the Prompt Journey
- Trailhead — The Agentforce Trust Layer: Follow the Response Journey
- Salesforce Help — Grounding with related list merge fields
- Salesforce Developers Blog — Ground your prompt templates with data using Flow or Apex
- Salesforce Developers — Access Models API with REST
- Salesforce Developers — Specify languages and locales with Models API
- Salesforce Developers — Data masking with Models API
- Salesforce Developers — Toxicity scoring with Models API
- Salesforce Help — Einstein Trust Layer: prompt journey
- Salesforce Help — Einstein Trust Layer: response journey
- Salesforce Help — Data masking limitations in Agentforce
- Salesforce Help — Large language model data masking
- Salesforce Help — Select what data to mask
- Salesforce Help — Einstein Trust Layer region and language support
- Salesforce Help — Prompt injection detection (Beta)
- Salesforce Help — Review toxicity scores
- Salesforce Help — Audit and feedback data model
- Salesforce Developers — Supported models
- Apex Reference — ConnectApi.EinsteinLLM
- Apex Reference — EinsteinLlmGenerationSafetyScoreOutput
- Salesforce — Agentforce privacy FAQ (PDF)

