Probably has raised $9 million in seed funding from Andreessen Horowitz to build a validation harness that aims to push large language model accuracy to 99.99%, the kind of reliability common in deterministic software but rare in AI. Founder Peter Elias is betting that hallucinations can be engineered out at the system level rather than waited out at the model level. The pitch lands at a moment when token costs are climbing and customers are openly trimming AI budgets.
The first product is a data science tool that returns quick answers from complex datasets, each accompanied by a citation and an audit trail. Behind the user-facing simplicity sits what Elias calls a data science mech suit: the LLM's first-pass answers run through a deterministic validator, which rejects any result that fails to match the underlying dataset. The model itself has been trained against that validator, so the whole loop is tuned for speed and correctness rather than open-ended generation.
The architectural payoff is what makes the round interesting. Probably's current system runs on a model four classes weaker than the frontier, light enough to execute on a desktop rather than a data center. That collapses a large chunk of the token bill that comes with calling GPT-class or Claude-class APIs for every query, and it sidesteps the data-residency headaches that keep regulated industries away from cloud inference.
Key facts
- 01Probably raised $9M in seed funding from Andreessen Horowitz to build a harness system that targets 99.99% LLM accuracy.
- 02The current product runs on a model four classes weaker than frontier models, enabling local hardware deployment instead of data centers.
- 03First product is a data science tool that returns answers with citations and an audit trail, validated against a deterministic checker.
- 04Founder Peter Elias plans to extend the engine to accounting, medical services, and other precision-sensitive use cases.
Elias frames the design philosophy bluntly: the harness does the work the model used to be asked to do.
The economics here are the story. Frontier model pricing has been moving the wrong way for buyers, with several labs raising per-token rates as reasoning modes and longer context windows roll out. A system that delivers higher accuracy on a smaller, locally hosted model inverts the usual AI cost curve, where every quality gain has historically meant a bigger bill.
Elias says the same engine extends well beyond data analysis. Accounting and medical services are on the roadmap, and the broader target is any precision-sensitive use case where a wrong answer carries real downside. Those are exactly the verticals where enterprise buyers have been most hesitant to deploy generative AI in production, and where deterministic guarantees matter more than conversational range.
The competitive read is that Probably is attacking a problem the major labs have not seriously pursued.
That observation has teeth. OpenAI, Anthropic, and Google all monetize on token volume, and a system optimized to need fewer correction rounds, on smaller models, is structurally at odds with their business model. It is a familiar pattern in enterprise software: the incumbents optimize for capability and consumption, while a startup builds the boring scaffolding that makes the capability usable in regulated environments.
The skeptical case is straightforward. A harness-plus-validator architecture works only when the underlying domain can be checked deterministically — a dataset query, a calculation, a structured lookup. Open-ended reasoning tasks, where there is no ground truth to validate against, are harder to fit into this mold. Probably's roadmap into accounting and medicine is plausible precisely because those domains have rules; expansion into less structured territory is a different engineering problem.
Probably is a small bet on a contrarian thesis: that reliability, not raw capability, is the gating constraint on enterprise AI adoption, and that the path to reliability runs through harness engineering rather than bigger models. If the thesis holds, the model layer becomes more commoditized than the labs would prefer, and value migrates to the orchestration and validation layer sitting above it. A $9M seed from a16z is a cheap option on that shift, and it is the kind of structural argument that gets more interesting as inference budgets tighten.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




