What agentic AI in private markets cannot do — and the cost of finding out.
Every fund is being pitched an agentic AI investing tool right now. Most of them demo well. Most of them will not survive eighteen months in your real book. Here is why, with the numbers.
The category is loud right now. New entrants every week. Investor decks promise an “AI investment associate” that reads the data room, builds the model, drafts the memo, raises the flags. The demo runs in four minutes. The room is impressed.
Then the deal closes, the agent goes into production, and the work starts to look different. Numbers that look right are wrong in ways your team finds in week six. Add-backs that should have been flagged sail through. Customer concentration is computed against the wrong subject. A portfolio DSCR is averaged when it should have been summed. The output is plausible enough that nobody catches it until a covenant test or a quarterly close forces a reconciliation.
This is not a prompt-engineering problem. It is the structural failure mode of probabilistic systems applied to evidence-grade investment work. Below is what you should expect, evaluated honestly.
The success rate looks fine in demo. It collapses in production.
On a curated demo deal, agentic-AI investing tools land 85–95% of the answers correctly. The deal team rehearses the prompts; the documents are clean; the questions are scoped.
On a real customer's first deployment — a deal the team didn't pick, with the documents the borrower actually sent — success rates drop to 30–50%. The variance is largest where lanes mix: a deal that's half-corp, half-real-asset, with multiple legal entities in the stack and a multi-quarter reconciliation, breaks the agent in ways the demo never tested.
By month eighteen at an institutional-grade customer — an LP, a regulated lender, anything that touches an audit — under 20% of these deployments will still be in their original shape. Most will have been quietly downgraded to “co-pilot for the analyst.” Which is to say: the human checks everything, and the productivity claim quietly disappears.
Time to ship a demo: months. Time to be right at scale: years.
A reasonably staffed team can ship an agentic loop with RAG over the data room in 4–8 months. The demo will look 80% done. It is actually closer to 15% done. The remaining 85% is the deterministic infrastructure underneath — the layer that catches all the silent miscategorizations the agent makes.
Time to pass a customer's first regression test (a known-bad deal where the right answer must surface a specific blocker): 18–30 months. During this window the team will discover their conservation-arithmetic bugs (averaged DSCRs, netted gaps), their subject-identity bugs (rent roll matched to wrong property, two LLCs with the same sponsor conflated), and their lane-shape bugs (RE shocks applied to a corp deal). Each discovery requires building part of the substrate the LLM-first product skipped on day one.
Time to ship a system that holds up at scale, across deal lanes, across audit cycles: 5–8 years for a well-funded team that knows what they are trying to build. Most teams will not get there because they will keep doubling down on bigger models instead of building the deterministic layer below. Every quarter that passes, the silent-wrongness surface area grows faster than the verification layer.
The token economics are quietly large.
For a single agentic-AI pass over a real deal — read the documents, extract KPIs, run covenants, draft a memo, raise flags — expect 200K–500K input tokens and 100K–300K output tokens. At current API rates, that is $5–15 per pass.
But real deals get re-analyzed. Quarterly cycles, evidence updates, continuous monitoring — typical fund operations run 10–60 passes per deal per year, which means $250–1,500 per deal annually before retries.
At scale:
A mid-size PE fund with 30 portfolio companies will burn $7K–45K/year in raw inference. A direct-lender with 200 borrowers in monitoring will burn $50K–300K/year. A multi-strategy alts manager with 1,000 deals will burn $250K–1.5M/year.
Five things prompt engineering will not fix.
Conservation arithmetic. Portfolio DSCR is the sum of NOI over the sum of debt service — not the average of per-property DSCRs. Refinance shortfalls are summed across properties, not netted. Liquidation gaps are summed even when other properties have surplus. Both forms of the answer are linguistically plausible. An LLM that has not been wrapped by a deterministic policy layer will at some point compute the wrong one — and the customer will not catch it.
Subject-identity guarantees. Multi-document deal rooms have multiple subjects: properties, legal entities, sponsors, customers. Without a deterministic guard layer that fails closed, an agent will at some point match a tax return to the wrong property than the rent roll covered, conflate two LLCs that share a sponsor, or route a market comp to the wrong subject. Each of these errors looks correct on its face. They surface as losses, not as exceptions.
Lane-shape correctness. A corporate private-credit deal needs corporate stress packs and covenant axes. A real-estate asset needs DSCR and rent-roll analytics. A hybrid needs both, without blending them. An LLM without lane-aware projection will silently apply RE-biased Monte Carlo to a corp deal, surface rent-roll KPIs where there is no rent, or generate cap-rate analytics on a private-credit refi. Customers rarely notice; they trust the wrong number.
Cross-source reconciliation that stays transparent. Management pack says revenue X. Audited financials say Y. Q of E normalizes to Z. A useful platform shows all three with reasoning, picks one explicitly, and logs the conflict. An LLM-only agent picks one silently or hallucinates a synthesis. The customer cannot audit the choice, which means a year later nobody can defend it. This is why Capital Refinery ships a document-review surface instead of a synthesis: when the CIM says 42 units and the rent roll enumerates 38, or a workbook arrives password-protected, the disagreement becomes a review item with both values as one-click choices — and the resolution (who chose what, and why) is written to an audit log, not lost in a chat transcript.
Saying “I don't know.” Confabulation is the LLM's default mode. A useful platform for IC and IR work has the opposite default: it tells you what is missing, what would sharpen the call, and which signal it is deliberately suppressing because the data is not yet decision-ready. That discipline cannot be added with a prompt. It is an architectural commitment that has to be made on day one. Most agentic products did not make it.
How to evaluate the difference, in fifteen minutes.
Three questions to ask the vendor in your next pitch:
1. Show me a downloadable deal pack where every KPI cites its source cell. Not a screenshot of a UI claiming traceability. An actual zip file you can take with you and inspect on your own laptop. If they cannot produce one, the traceability claim is marketing, not infrastructure.
2. Show me the same deal run twice. Are the dossiers byte-identical? If the platform's own output is not deterministic on identical evidence, no fingerprint, anchor, or share-link can mean what the marketing claims. Reproducibility is the floor.
3. Ask the platform to recommend a deal where it can't. A good evidence platform will tell you which KPIs are missing, which evidence would sharpen the decision, and refuse to render an IC-grade memo against insufficient data. An LLM-first product will produce a memo regardless. That difference is the entire moat.
If you are evaluating an agentic-AI investing tool right now, download our deal packs and run the same three tests against the alternative. The difference is visible in the file structure before you open the first PDF.
Don't trust the dashboard. Trace the numbers yourself.
Three downloadable deal packs across the lifecycle. Every KPI cites its source cell. Open kpis.csv, follow the citation, the number matches. No app, no login required.
