Koa is based on Nvidia Nemotron and trained with synthetic business tasks. Salesforce says it makes fewer CRM action errors, but the benchmark and the customer pilots still need outside study.
New to this? Read it in simple words
- Salesforce and Nvidia introduced Koa, a reasoning model for sales, service, and other customer work. It is based on Nvidia Nemotron and trained with invented business tasks.
- Salesforce says Koa matches or beats leading models on its own CRM benchmark. It also says Koa makes three times fewer action errors.
- Such errors matter, because a wrong record update can harm a customer.
- Salesforce made the benchmark and reported the result, so outside tests are still needed. For now, only selected customers can test Koa.
- Reasoning model
- An AI model that works through a problem step by step before answering.
- Benchmark
- A standard test used to compare AI models.
- CRM
- Short for customer relationship management: software that companies use to track customers and sales.
Koa is a specialised model, not a new general chatbot
Salesforce built Koa by adding new training to Nvidia Nemotron 3 Super. The goal is to make the model reason through several business steps. Examples include updating a sales opportunity, routing a service case, or scheduling a follow-up.
The training used synthetic scenarios instead of real customer data. Salesforce says the scenarios covered more than 14 industries. Each one joined a business role, a goal, and the tool calls needed to finish the work.
This narrow training can be useful. A smaller specialist may understand company actions better than a general model. It can also run inside a controlled service, with clearer rules about data and model access.
Salesforce reports fewer action errors on its own benchmark. Independent results and real customer failures still need to be published.
The main performance claim needs outside testing
Salesforce says Koa matches or beats leading models on its CRM benchmark and makes three times fewer action errors. The test includes common business tasks. However, Salesforce created the benchmark and reported the result.
A strong test should publish the task set, scoring rules, model versions, tool setup, and failed examples. Independent teams should also repeat the work. Without these details, readers cannot compare the result fairly with another model.
Tool errors matter more than a polished answer. A wrong field update can harm a customer or change a forecast. Tests should measure bad actions, safe refusal, recovery, and whether the model asks for help when information is unclear.
The first release is a pilot, not wide use
Koa is available to selected Agentforce pilot customers. Salesforce expects a wider US release in winter 2026. It says several companies are testing the model, including Xero, Formula 1, and UChicago Medicine.
Salesforce controls the model weights and runs training and use inside its own trust boundary. That may help customers with strict data rules. It does not remove the need for access limits, human approval, logs, and incident reporting.
The important question is simple: does the model complete real work with fewer harmful errors? Pilot users should publish clear task results and limits. A company benchmark is a useful start, but it is not the final answer.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1Primary source · September 15, 2026Announcing Koa: Salesforce’s first CRM reasoning model, built on Nvidia Nemotron Salesforce
- 2Primary source · September 15, 2026Salesforce unveils AIforce, bringing the full power of its platform to any interface Salesforce
- 3Research · September 15, 2026Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear Yahoo Tech / TechCrunch
We used the joint announcement for the model design, benchmark claim, pilots, and release plan. We used Salesforce’s wider AIforce note and independent TechCrunch reporting for context. We label the benchmark result as a Salesforce claim.