Audit our thinking before you audit our output.
Capital Refinery is an institutional measurement layer, not a consultancy and not an AI scoring product. The credibility of the artifact comes from the discipline of the engine that produces it. This page consolidates the doctrine: five principles, ten axes, two gates, the arbitration logic, the Lift Ledger, the refusals, and how any output is independently verifiable. If a section here doesn't match what runs in production, file a bug.
Five principles
The methodology rests on five doctrine principles. They are not editorial preferences. They are structural constraints on what the engine is allowed to claim — enforced in the grading code, tested as refusals, preserved across every assessment.
- Principle 01Deterministic-first
The brain is deterministic extraction and arbitration. Local LLM is thin adjudication where rule-based logic genuinely cannot decide — never grading, never narrative, never policy. Every promoted KPI carries candidate_id, source_ref, and method. If a deterministic path can produce the answer, the deterministic path wins.
- Principle 02Documents win
When operator-claimed inputs and document-derived inputs disagree, documents win. Operator narrative can calibrate the band by one notch when document evidence already supports the claim. Conflicts (claim says X, docs show Y) surface as named blockers, never silently averaged.
- Principle 03not_observable, never imputed
When evidence is thin, the engine refuses to grade and marks the sub-axis not_observable with a reason. It does not impute zero, default, or median. Worse-looking on the artifact, harder to fake, and the only posture that survives sophisticated review.
- Principle 04Observation, not attribution
The Lift Ledger measures what changed between two evidence-backed snapshots. Causality belongs to whoever did the work, never to the engine, never to the partner who hired the engine. 'EBITDA moved from $8.0M to $12.5M' is provable. 'Our work caused 18% growth' is not.
- Principle 05Refused inputs
Sentiment, polish, presentation tone, narrative coherence, management warmth, and 'AI maturity vibes' are refused as graded inputs. Only documents and elapsed-time signals count. LLM-derived scores on human prose are opinion in numeric clothing — and the engine refuses to dress opinion up as evidence.
The Institutional Readiness Assessment: ten axes
Every IRA grades a deal across ten universal axes. Sub-lane overlays (corp.distribution, corp.dentistry, re.industrial, re.multifamily, etc.) adjust thresholds — never the axis structure. New industries do not get their own grading code; they get tighter or looser bands on the same engine.
- Financial consistencyGate
- Data integrityGate
- Reporting maturityStd
- KPI completenessStd
- Operational riskStd
- Stress toleranceStd
- GovernanceNew
- Management responsivenessNew
- Key-person dependencyNew
- Customer concentrationNew
The Financial Truth Engine
Underneath the axis grades is a six-layer arbitration system that produces the canonical KPI values the IRA reads from. The layers exist because financial truth in middle-market deals is almost never single-sourced — the same revenue line shows up in a CIM, a tax return, a QoE report, a model, and a management presentation, and they often disagree.
- 1. Parsers. Per-document- shape extractors (rent rolls, T-12s, AR aging, credit agreements, QoE reports, etc.). Each emits structured candidates with provenance back to the page.
- 2. Arbitration. When multiple candidates exist for the same KPI, the engine picks a winner by method-priority, plausibility floor, symmetric median-distance outlier penalty, and cross-KPI consistency sweep. Cross-KPI consistency surfaces tension; it does not auto-demote winners. The advisory layer flags; the canonical layer stays clean.
- 3. Construction health. Validates that the underlying KPI math holds (revenue ≥ 0, DSCR derives correctly from EBITDA / debt service, unit-normalized values within plausible bands per sub-lane).
- 4. Reconciler. Logs every conflict with both losing candidates so the trail is inspectable. Critical-KPI conflicts (DSCR, FCCR, revenue, EBITDA, total debt) become named blockers when unreconciled.
- 5. Dependency engine. Cross-KPI dependencies: leverage requires total debt and EBITDA, FCCR requires fixed charges and EBITDA. If a dependency is missing or contested, the derived KPI stays
not_observable. - 6. Decision gate. The engine reports
is_decisionable— a boolean answering whether the evidence trail is complete enough to promote IC-grade output. The platform refuses to render IC-bound artifacts when this is false.
The whole system is fail-closed by design. The platform’s honest default is “not yet decisionable,” not “here’s our best guess.”
The Lift Ledger: four layers
When an IRA is re-run after work has landed (the “Re-IRA delta”), the engine produces a Lift Ledger — a structured comparison between two evidence-backed snapshots, organized into four measurement categories. The Lift Ledger is the artifact that turns Capital Refinery from a one-shot review into a measurement layer with a renewal cadence.
Observed post-engagement movement. The Lift Ledger is an evidence-backed measurement of state change; it is not an attribution opinion about which interventions caused which changes. Causality belongs to the operator, the consultant, and the underlying business — the Ledger measures only what moved between snapshots, where, and how we know.
Each layer has its own “not observable” semantics. If a canonical KPI exists on one snapshot but not the other, the Economic Lift line for that KPI reads not_observable: promoted_kpi_missing_on_one_snapshot rather than imputing a zero or carrying the prior value forward. The Process Lift layer reads management responsiveness as throughput (median response days, overdue request count), never as quality / tone / sophistication.
Verification: every artifact is independently checkable
Every IRA snapshot carries a deterministic fingerprint computed from its evidence trail and a public, HMAC-signed verification token. The token resolves to a stripped public view at /p/ira/<token>. A buyer, lender, or board reviewer can:
- Verify the artifact came out of the engine (not a Word document edited after the fact) by comparing the fingerprint on the docx / xlsx export against the fingerprint at the verification URL.
- See the same composite verdict, axis bands, and named blockers a sophisticated reviewer would inspect — minus operator-internal PII.
- Read the cause label attribution: the snapshot tells you why it was generated (ingestion / decision-commit / quarterly recompute) and which source document drove the change.
Exports also embed an IC-anchor fingerprint when the snapshot was committed to a decision. The verification page renders an Export Alignment Banner: Aligned / Superseded / No Current Anchor. A signed export that gets superseded by new evidence remains verifiable but is correctly flagged as out-of-date.
What we refuse to grade
The methodology is defined as much by what it refuses as by what it produces. Every refusal below is structural — not a stylistic choice — and each one exists because including it would erode the institutional weight of the artifact.
- No AI maturity score. “Digital transformation readiness” and “AI maturity” are vibes wearing a number. Not produced.
- No consultant ROI claim. The engine measures change between two snapshots. It does not claim a partner caused the improvement, and partners using the artifact cannot use it that way under our partner-as-agent legal frame.
- No imputation. Missing data means
not_observablewith a reason. Never zero, never default, never median. - No narrative grading. Sentiment, tone, presentation polish, management warmth, decision-record sophistication, and communication style are not graded. LLM-derived sentiment scores on human prose are opinion in numeric clothing.
- No black-box bands. Every band on every axis derives from named, inspectable evidence. Every promoted KPI carries
candidate_id,source_ref, andmethod. The full reasoning tree is exposed via the Why-this-number overlay. - No outcome guarantee. The Lift Ledger reads in observation language, not promise language.
- No customer theater. Cedarbrook is a methodology fixture, not a customer. The walkthrough discloses this directly. We do not show logos we have not earned.
Where to inspect this in production
Reading the methodology is the first step. Watching it run on a real deal is the second. Two proof cases — one for each audience — and three customer-facing surfaces.
Two proof cases, by audience
- Falcon Services — the buy-side proof case. A real PE/PC services deal under watch posture. The IC dossier, per-cell evidence trail, stress lab, covenant forecast, decision lifecycle, and verification artifact — rendered live across /watch, /sample-falcon (the seven-artifact subscriber-tier buy-side proof pack), and /sample-diagnostic (the full output of the same-day /diagnostic product on this deal — a different scope, not a lighter version of the proof pack). PE, credit, and family-office investment teams should start here.
- Cedarbrook Foods — the sell-side proof case. A real corp-PC distribution deal preparing for sale. Composite verdict, axis grid, two drill-ins, the Re-IRA delta, and the public verification page — with the seller-readiness pack (IRA assessment + IC memo + executive report + workbook + KPI summary) downloadable at /sample-evaluation. Owners, advisors, fractional CFOs, and modernization consultants should start here.
Three customer-facing surfaces
- Pressure-test one deal. Buy-side. Send a CIM, a credit agreement, or a model. Same-day diagnostic on your laptop. The smallest scope that exercises the engine on a deal you actually care about.
- Readiness Gap Review. Sell-side. $4,500. Five business days. The smallest scope that exercises the engine on an operating company. Credit toward a full IRA within 60 days.
- Modernization Impact Review. Partner-channel. Independent measurement layer attached to AI / automation / fractional-CFO engagements. $3,500 baseline + $2,500 Re-IRA delta. Partner-as-agent legal frame.
The methodology is the same in all three. The packaging differs by audience.
Inspect the methodology against a real deal.
The fastest way to evaluate whether this discipline holds is to run it on something you actually care about. Bring a deal, a portfolio company, or a business you're preparing to sell.

Built by Duaine McDonald, an AI strategist and fractional Chief AI Officer with 20+ years across enterprise automation, transformation, governance, and AI strategy. The consulting practice runs separately as Enterprise Refinery; Capital Refinery is the software layer that came out of those conversations — the institutional measurement artifact operating companies, advisors, and investment teams kept asking for and could not find anywhere else. Early engagements are reviewed directly by the founder.
Capital Refinery is early. We do not show customer logos we have not earned.
Instead, we show the methodology, two fixture-based proof cases (one per audience), the verification flow, and the live artifacts a buyer, lender, board reviewer, or advisor would inspect. Most early-stage products would invent a logo wall. We refuse on purpose — the same discipline that makes the artifact credible.