Arga Labs raised a $10 million seed round to build sandbox training environments for enterprise AI agents, the company announced Wednesday. General Catalyst led the round, with Box Group, Emergence, Gradient, and SV Angel participating. The pitch is that enterprise agents keep failing in production because the tools used to train them don't resemble the software they're deployed against.
Most testing environments for agents settle for a stateless API endpoint — a stripped-down mock that responds to calls but doesn't preserve state between runs. Arga Labs instead builds full-scale digital twins of programs like Salesforce, Workday, and Outlook, cloning the entire application with its permission systems and web hooks intact. The result is an environment an agent can be dropped into, fail inside, and be reset from, thousands of times over.
CEO and co-founder Philip Li illustrates the core problem with a mundane sales scenario. A prospective client creates a lead in Salesforce while a colleague reaches out separately through HubSpot, and an agent is supposed to reconcile the two.
Key facts
- 01Arga Labs raised a $10M seed round led by General Catalyst, with Box Group, Emergence, Gradient, and SV Angel participating.
- 02The startup builds full-scale digital twins of enterprise software including Salesforce, Workday, and Outlook — not stateless API mocks.
- 03The twins preserve permission systems and web hooks, letting agents be reset and retrained across tens of thousands of runs.
- 04Co-founder Philip Li frames the gap as the same reinforcement-learning tooling that made AI coding tools competitive, now applied to business software.
That kind of cross-system ambiguity is where current agents fall over. Reinforcement learning is the standard fix — run the scenario tens of thousands of times, keep the strategies that work, discard the ones that don't. But you can't do that against a live Salesforce tenant. There's no clean way to reset a production CRM, and no practical way to clone one at the scale RL requires.
Arga Labs' twins are designed to be resettable and clonable by default. Because the company controls the environment end to end, it can spin up many parallel instances, wire them together to simulate a full employee workspace, and reset the whole assembly between runs. Agents can be trained on the interactions between programs — a lead in one system, a calendar in another, an inbox in a third — rather than on isolated API calls.
Li frames the market opportunity as closing a reinforcement gap. AI coding tools advanced quickly in part because software engineering already had sophisticated infrastructure for deploying, reversing, and analyzing code changes. That tooling made it straightforward to build RL environments for coding tasks, and coding agents improved accordingly. Business software has no equivalent, which is why agents operating inside Salesforce or Outlook still look primitive next to agents writing Python.
General Catalyst's Yuri Sagalov, who runs the firm's seed program, said the need for agentic testing infrastructure is growing as more enterprise budget shifts toward agents rather than seats.
The competitive frame here is broader than Arga. Enterprise agent evaluation has become one of the most active seed-stage categories of 2026, alongside earlier bets like QueryStory's $6M raise to make enterprise AI answers auditable. Every large model provider — Anthropic, OpenAI, and the frontier labs behind agent frameworks — needs the same thing Arga is selling: a way to prove that an agent won't break a customer's CRM before the customer lets it near production data.
Whether Arga's digital-twin approach wins depends on how faithfully the twins track the real applications. Salesforce and Workday ship changes constantly, and a twin that drifts from the live product becomes a training environment for the wrong software. The company will have to invest in keeping its clones current, and it will have to convince buyers that agents trained in a sandbox generalize to the messier live systems on the other side.
The bet worth watching is whether enterprise agents follow the coding-tools curve. If Arga and its peers succeed in building the RL infrastructure that business software has never had, the same compounding improvements that turned coding assistants into serious engineering tools should show up in sales, finance, and operations agents within the next two years. That would reprice a lot of enterprise software, and it would make the sandbox layer — currently invisible to end users — one of the more valuable pieces of the agent stack.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




