an ai product sold to regulated institutions has to prove it is right. i built the measurement, the frameworks and the scoring that let it.
evaluation & qa program
golden datasets, llm-judge scoring, daily human review. relevancy accuracy doubled.
due-diligence library
frameworks across defi, stablecoins, rwas and more. 60 to 100+ criteria each.
deterministic risk rating
seven domains, ~50 evidence-gated flags, one 0–100 score a regulator can re-derive.
financial-crime research
how the first line investigates, and how the second line tests it.
the daily evaluation loop, measured weekly