← Back to work

AI Review: Evaluation Mode

Trustworthy AI for compliance-grade investment due diligence: I led design and product strategy for an AI evaluation system that rates, flags, and reasons across due-diligence responses a deterministic rules engine could never touch, all while keeping human judgment in the loop through in-context override and a permanent audit trail.

AI Review: Evaluation Mode
Timeline
1-2 months
My Role
Lead Designer + Product Strategy
Team
Engineering, Product, QA
Platform / Tools
B2B SaaS · Due Diligence · Figma
Goal

Replace a manual, deterministic rules engine with an AI evaluation system that compliance-driven investment teams could actually trust: one that could rate, flag, and reason across the response types the old system couldn't, attachments, grids, multi-line inputs, without ever taking human judgment out of the loop.

Challenge

The review layer automated the easy 20% and left the risky 80% to human effort.

The engine ran on deterministic if/then rules, applied by hand to text responses. It worked for simple fields, but the response types that carried the most due-diligence risk (attachments, grids, multi-line inputs) fell outside what rules could evaluate at all. So analysts either skipped structured review or did it manually, off-platform.

At scale, that's not a UX annoyance. In a compliance context, a missed flag or an undocumented override is an audit risk. And the churn told me the failure was already priced in: a handful of power users were quietly holding the system together, and everyone else had walked.

Discovery

I read the system before I redesigned it, and the data made the first decision for me.

Digging into analytics, I found the Advanced rules tab, one of three modes in the rule-creation modal, had produced 6 rules in two years. That wasn’t a discoverability problem. Nobody needed it. So before adding anything, I made the case to remove it and simplify the creation surface. Subtract, then build.

Client conversations sharpened the churn diagnosis: analysts wanted AI-assisted review, had seen it elsewhere, and didn’t believe a manual system could get them there. And the timing was on my side. The platform was already investing in AI for its Autofill feature, so the infrastructure and the organizational confidence existed. The investor-side review workflow was the right next bet, and I could argue why now, not just why.

Half my week went to manually checking grids and attachments the rules engine couldn't even read. If something slipped through, that was on me, not the tool.
Analyst · institutional due diligence

Analysts weren’t short on diligence, they were short on a tool that could evaluate the formats that carried the most risk. The gap wasn’t effort. It was coverage.

3
response types the old engine couldn't touch
80%
of review effort left to manual work
My approach

I designed AI Review to coexist with the rules system, not replace it. Analysts can mix Standard rules and AI Mode within the same template, so the trust they’d already built in what they understood stayed intact.

The bigger strategic bet was trust architecture. Introducing AI into a compliance workflow without override, confidence, and audit isn’t a feature, it’s a liability. So I scoped those in from v1 and framed them to stakeholders as client-retention infrastructure, not feature overhead. That framing is what got the backend investment approved.

The build ran in four phases: scoping the architecture with product and engineering, rebuilding the existing flag, rate, and override flows so override sat inside the review hierarchy, consolidating rules and prompts into the template builder, and a rigorous pilot in a controlled environment before any production rollout.

Impact & outcomes
305
AI rules created across 7 templates
11
Projects evaluated

Beyond the numbers: analysts who had disengaged from the manual system engaged actively with the AI interface in demo, and stakeholders positioned it as a competitive differentiator against peer platforms still running deterministic review.

01

Rule creation in the template builder: deterministic rules, ruled out

I cut the Advanced tab the data had already killed, then rebuilt creation around two modes: Standard (flag / rate) and AI Mode (natural-language prompts). Then I pulled every question’s rules and AI prompts into a single panel in the template builder, sitting alongside the Autofill prompts already living there. One place to see everything applied to a question, instead of hunting across a separate rules page.

Configure Prompts panel with Set up AI Automation options for Auto-Fill, Review, and Insight prompts, next to the Create Review Prompt modal showing Standard and AI Mode, applied to a specific question, with rating type, workflow, and prompt name fields
02

Project-level AI summary

Once a template is attached to a project, evaluation runs and rolls up into a project-level summary: rating, confidence, flags, and completeness across the whole review. The analyst reads the shape of the project before drilling into a single response.

Evaluation Summary modal: 75 of 100 responses evaluated, an AI-written summary of what needs attention, and an Evaluation Results table breaking down flag, rating, incomplete, and passed outcomes by count and confidence
03

AI evaluation at the question level

The same evaluation, scoped to the individual response. I deliberately confined AI evaluation to project and question level only, not every possible surface, to keep the mental model simple and the first release trustworthy. Each response surfaces an AI rating, a confidence score (High / Limited), flag status, and expandable reasoning showing the exact Assessment Criteria the AI evaluated against. The confidence label isn’t decoration, it shapes behaviour: High means less manual checking, Limited is a prompt to look closer.

Question-level AI evaluation: a rating with limited confidence, expandable AI reasoning, and the assessment criteria the AI evaluated against, alongside the flag status
Question-level flag and completeness evaluation: an issues-found flag with its assessment criteria, and a completeness rating listing which required fields passed and which are still incomplete
04

Override the AI, or fix the prompt

When the AI is wrong, the analyst has two moves, not one. Override the rating in place, logged with name, timestamp, original AI result, and reason. Or, if the prompt itself is off, edit the prompt so the correction holds for every response after it. I designed override before anyone required it, because a prompt-driven system fails in two directions, and the person closest to the prompt needs to correct both.

Override modals for flag and rating: overriding the AI review unflags or reflags a response, or changes its rating, with a required reason for override and a note that the override is logged for audit
Edit Prompt modal showing the named prompt text, an option to re-run for the selected question or re-run and update it across the entire project and template, and advanced options including marking it a universal prompt
05

Review history as audit, next to response history

I placed review history directly beside the existing response history at the question level, so every AI result, every override, and every reason lives in the response’s own timeline. Accountability isn’t buried in a log somewhere; it’s loud, in context, exactly where the next analyst or auditor is already looking.

Questionnaire response panel next to an Audit Trails panel showing Review History: current version overridden by a manager, with the original AI rating and reasoning, the new flag status, the override reason, and an option to reveal the AI's original reasoning
Portfolio highlights
  • Led design and product strategy for a 0-to-1 AI evaluation system in a compliance-grade fintech platform.
  • Used two years of usage data to remove the Advanced tab before building, simplifying the surface with zero stakeholder pushback.
  • Designed override and audit before either was required, anticipating how prompt-driven AI actually fails.
  • Resolved the core conflict between AI override and the existing flag/rate/exclude system through a coherent modal rebuild, not bolt-ons.
  • Won backend scope approval by framing trust architecture as retention infrastructure, not feature cost.
  • Made the strategic call to ship without the summary modal, and documented it as the next-phase priority.
Next project · Breathefree App

Let's build something trustworthy.

Get in touch →