Trustworthy AI for compliance-grade investment due diligence: I led design and product strategy for an AI evaluation system that rates, flags, and reasons across due-diligence responses a deterministic rules engine could never touch, all while keeping human judgment in the loop through in-context override and a permanent audit trail.
Replace a manual, deterministic rules engine with an AI evaluation system that compliance-driven investment teams could actually trust: one that could rate, flag, and reason across the response types the old system couldn't, attachments, grids, multi-line inputs, without ever taking human judgment out of the loop.
The review layer automated the easy 20% and left the risky 80% to human effort.
The engine ran on deterministic if/then rules, applied by hand to text responses. It worked for simple fields, but the response types that carried the most due-diligence risk (attachments, grids, multi-line inputs) fell outside what rules could evaluate at all. So analysts either skipped structured review or did it manually, off-platform.
At scale, that's not a UX annoyance. In a compliance context, a missed flag or an undocumented override is an audit risk. And the churn told me the failure was already priced in: a handful of power users were quietly holding the system together, and everyone else had walked.
I read the system before I redesigned it, and the data made the first decision for me.
Digging into analytics, I found the Advanced rules tab, one of three modes in the rule-creation modal, had produced 6 rules in two years. That wasn’t a discoverability problem. Nobody needed it. So before adding anything, I made the case to remove it and simplify the creation surface. Subtract, then build.
Client conversations sharpened the churn diagnosis: analysts wanted AI-assisted review, had seen it elsewhere, and didn’t believe a manual system could get them there. And the timing was on my side. The platform was already investing in AI for its Autofill feature, so the infrastructure and the organizational confidence existed. The investor-side review workflow was the right next bet, and I could argue why now, not just why.
Half my week went to manually checking grids and attachments the rules engine couldn't even read. If something slipped through, that was on me, not the tool.
Analysts weren’t short on diligence, they were short on a tool that could evaluate the formats that carried the most risk. The gap wasn’t effort. It was coverage.
I designed AI Review to coexist with the rules system, not replace it. Analysts can mix Standard rules and AI Mode within the same template, so the trust they’d already built in what they understood stayed intact.
The bigger strategic bet was trust architecture. Introducing AI into a compliance workflow without override, confidence, and audit isn’t a feature, it’s a liability. So I scoped those in from v1 and framed them to stakeholders as client-retention infrastructure, not feature overhead. That framing is what got the backend investment approved.
The build ran in four phases: scoping the architecture with product and engineering, rebuilding the existing flag, rate, and override flows so override sat inside the review hierarchy, consolidating rules and prompts into the template builder, and a rigorous pilot in a controlled environment before any production rollout.
Beyond the numbers: analysts who had disengaged from the manual system engaged actively with the AI interface in demo, and stakeholders positioned it as a competitive differentiator against peer platforms still running deterministic review.
I cut the Advanced tab the data had already killed, then rebuilt creation around two modes: Standard (flag / rate) and AI Mode (natural-language prompts). Then I pulled every question’s rules and AI prompts into a single panel in the template builder, sitting alongside the Autofill prompts already living there. One place to see everything applied to a question, instead of hunting across a separate rules page.
Once a template is attached to a project, evaluation runs and rolls up into a project-level summary: rating, confidence, flags, and completeness across the whole review. The analyst reads the shape of the project before drilling into a single response.
The same evaluation, scoped to the individual response. I deliberately confined AI evaluation to project and question level only, not every possible surface, to keep the mental model simple and the first release trustworthy. Each response surfaces an AI rating, a confidence score (High / Limited), flag status, and expandable reasoning showing the exact Assessment Criteria the AI evaluated against. The confidence label isn’t decoration, it shapes behaviour: High means less manual checking, Limited is a prompt to look closer.
When the AI is wrong, the analyst has two moves, not one. Override the rating in place, logged with name, timestamp, original AI result, and reason. Or, if the prompt itself is off, edit the prompt so the correction holds for every response after it. I designed override before anyone required it, because a prompt-driven system fails in two directions, and the person closest to the prompt needs to correct both.
I placed review history directly beside the existing response history at the question level, so every AI result, every override, and every reason lives in the response’s own timeline. Accountability isn’t buried in a log somewhere; it’s loud, in context, exactly where the next analyst or auditor is already looking.