Home  ›  Insights  ›  Jev versus the Decisions API for Salesforce
AI & ML

Jev versus OpenAI’s Decisions API for Salesforce case routing

Two decision models, two weeks apart, both aimed at the same job: routing a case in milliseconds instead of writing about it. What is real, what is priced, what is proven, and a test kit to settle it on your own cases.

The short answer

Two decision models now compete for the same job inside a Salesforce org: picking the queue, the priority, or the next step for a case, and returning that choice in a few hundred milliseconds instead of writing a paragraph about it. TypeSafe AI's Jev has been generally available since 21 September 2026. OpenAI announced its Decisions API, built on GPT-6 Luna, at DevDay on 29 September.

Our position for a Salesforce service team, today:

  • Pilot Jev now, in shadow mode. It has a public price, public documentation, and an early body of independent tests, good and bad.
  • Put the Decisions API on the watch list, not the roadmap. It is in limited preview with no published price and no public API reference. You cannot cost or govern what is not documented.
  • Keep Einstein Case Classification where you already have the history. A decision model earns its place where labels change often, history is thin, or the decision sits outside the Salesforce UI.
What this page is, and is not. Every performance figure below is either published by the vendor, published by an independent tester (linked), or our own arithmetic on list prices, and each is labelled as such. We have not yet run either model on our test set. The test set and harness are published below so you can run it before we do.

What each one actually is

Both models take structured state in and return a typed answer from a list you define, rather than generated text. That is the whole point: a case router does not need prose, it needs one of six queue names and a confidence value it can threshold on.

As of 1 Oct 2026TypeSafe JevOpenAI Decisions API
StatusGenerally available since 21 SepLimited preview since 29 Sep; broad release promised “in the coming days”
Public API referenceYes: POST /v1/systemone, Python and TypeScript SDKsNone published
List price$0.042 per million input tokens; output not billedNot published. Luna's token rates cannot be assumed to apply
Answer typesChoice (up to 255 options), Score (2 to 10 levels), Noul (yes/no probability)One choice from a list you supply; option limit not published
Confidence returnedYes, for Choice and ScoreNot documented
Latency0.35 s median per passage in one independent testAbout 150 ms, vendor-stated, against roughly 1.6 s for Luna through the standard API
Routes inDirect, Vercel AI Gateway, Cloudflare, OpenRouterOpenAI only

Sources: OrcaRouter's Decisions API review, the Awesome Jev reference, and Hatchworks on Jev pricing.

What the independent evidence says

Three findings matter more than the launch-week view counts.

1. How you ask decides the accuracy

An independent phishing benchmark found Jev scored 62.6% when asked one broad question, against 81.3% for Claude Haiku 4.5. Split into five atomic signals, Jev reached 95.0% against Haiku's 93.2%, at $0.038 per thousand emails (reported by Kanerika). For case routing, that means asking “is this about money?”, “is this a security event?” and “is something broken?” separately, then combining, rather than one six-way question.

2. It trades a little accuracy for a lot of cost

On Banking77, a standard intent-routing benchmark, one evaluation put Jev at 0.78 against GPT-5.6 Terra's 0.85, at roughly one-fiftieth of the cost (same source). For a first-pass router with a human or Einstein behind low-confidence cases, that is often the right trade. For a router with no fallback, it is not.

3. Not every demo is real

Builder.io's CEO publicly asked people to stop posting fake Jev demos on 19 September, and a reply reported a browser-use test failing (CellCog's dated tracker). TypeSafe's own headline gains compare Jev against an average of flagship models, built by its own team. Treat every number, including ours, as a hypothesis until it has run on your cases.

Cost at Salesforce volume

Our arithmetic, on list prices only. Assumptions: one million cases a month, 600 input tokens per case (subject, description, account tier, channel, queue definitions), and 10 output tokens where output is billed.

ModelInputOutputPer month
Jev600M × $0.042 = $25.20Not billed$25.20
GPT-6 Luna, standard API600M × $0.10 = $60.0010M × $0.50 = $5.00$65.00
GPT-6 Sol, standard API600M × $2.00 = $1,20010M × $10.00 = $100$1,300
OpenAI Decisions APIPrice not publishedUnknown

Two conclusions follow. Against a flagship model, a decision model is a fifty-fold saving. Against OpenAI's own cheap tier, Luna, the list-price gap is about 2.6 times, not a hundred. At this volume the model bill is small either way. The cost that matters is the cost of a misrouted case: the hour a security incident spends in the billing queue. Accuracy on your ambiguous cases decides the business case, not the token price.

Luna prices from OpenAI's model page; Sol from LiteLLM's launch note.

Where it fits in a Salesforce org

The pattern is the same whichever model wins:

  • Trigger. A record-triggered Flow or Platform Event fires when a case is created.
  • Call. An Apex callout through a Named Credential, or a small middleware service, sends the case state to the decision model. Keep the API key out of the org.
  • Decide. The model returns a queue and a confidence value.
  • Act on confidence. Above your threshold, write the queue back to the case. Below it, fall back to Einstein Case Classification or a human triage queue.
  • Log everything. Store the model's answer, its confidence, and the final human-confirmed queue on the case. That is your accuracy dashboard, and your evidence when someone asks why a case went where it went.

Run it in shadow mode first: the model writes its answer to a custom field and routes nothing. Two weeks of shadow data tells you your real accuracy on your real cases, at no operational risk.

Two governance questions to settle before production. Data retention: zero-retention terms with TypeSafe are negotiated directly, not a default (Appinventiv). And regulated fields: strip anything you would not send to an external API before the callout, whichever vendor you use.

The test kit, and what we have not measured yet

We have published a synthetic, labelled Salesforce case set and the harness to score any decision model against it. No client data is involved.

  • salesforce-case-routing-testset.csv: 240 synthetic cases across six queues (billing, technical support, account management, security, data integration, sales), each with account tier and channel. 24 are deliberately ambiguous, sitting on a real boundary such as a billing dispute caused by an integration failure, with the secondary queue recorded.
  • decision_routing_harness.py: runs the set through Jev and GPT-6 Luna, and reports accuracy, accuracy on the ambiguous subset, p50 and p95 latency, and estimated cost at list price.

The set is small and template-built: 106 distinct subjects across 240 rows. It is a smoke test for the routing pattern, not a substitute for a shadow run on your own case history.

Not yet measured by us: accuracy, latency and cost for either model on this set; the Decisions API at all, since it has no public reference; behaviour on long case descriptions and email threads; and non-English cases. We will add our results to this page when we have them, with the raw output file.

Frequently asked questions

Can Jev replace Einstein Case Classification?

Not where Einstein already works well. Einstein learns from your closed-case history and needs at least 400 closed cases to start. Jev needs no training history, which suits new queues, frequently changing labels, and decisions made outside the Salesforce UI.

Is OpenAI's Decisions API cheaper than Jev?

Nobody can say yet. OpenAI has not published a price for the Decisions API, and its underlying model's token rates cannot be assumed to apply. Jev lists $0.042 per million input tokens with output not billed.

How accurate is a decision model at routing support cases?

It depends heavily on how the question is framed. Independent tests show broad single questions underperform, while decomposing the decision into several narrow yes/no or choice questions closes most of the gap with larger models. Measure on your own cases in shadow mode before routing anything.

How do we call a decision model from Salesforce?

Use a record-triggered Flow or Platform Event, then an Apex callout through a Named Credential or a small middleware service. Write the answer and confidence back to the case, and send low-confidence cases to Einstein or a human triage queue.