More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Probably just pulled in $9 million from Andreessen Horowitz to tackle one of AI’s hardest problems: hallucinations and factual slip-ups in large language models. Founder Peter Elias aims for that 99.99% accuracy you see in rule-based systems, but with the flexibility of AI. To hit that mark, Probably built what Elias calls a “data science mech suit”—an elaborate harness where every model output gets compared against a deterministic validator tied directly to the source data. If the answer doesn’t perfectly match, it never reaches the user.
Their first product sits on top of complex datasets, delivering quick summaries complete with citations and a full audit trail. Behind the scenes, the LLM generates an initial answer, then the validator either approves it or sends it back. Over time, the model learns from those checks, cutting down on ambiguity. That engineering trade-off means they can run on a much smaller model—Elias says it’s “four classes weaker than the frontier models.” You could spin it up on your desktop instead of a GPU farm, slashing token bills.
Probably isn’t stopping at data science. Elias thinks the same setup will work for any task where precision matters: accounting, medical records, compliance. He points out that big AI labs haven’t bothered, since their business model depends on users correcting errors—and paying for each interaction. With token costs climbing and AI budgets under pressure, a system that avoids mistakes and shrinks your compute bill could land well.
Questions about this article
No questions yet.