RAI is an AI due-diligence system for utility-scale solar. Upload a project’s documents, and nine specialist agents cross-examine them and return a readiness score with every finding cited. I designed the product and built the frontend: every screen an analyst touches.
Before anyone puts money into a solar farm, an investment team reviews the project’s paperwork: hundreds of pages written by different companies, from the cost estimate to the contracts to sell the power. Each document looks fine on its own.
The risk is when two of them disagree. If nobody catches it, the decision to invest gets made on the wrong number.
to build the whole project
for the equipment alone
The investors’ expected return is calculated on the lower number. At the quoted price, it drops below what the fund requires.
of power sold to buyers
of that power
The revenue plan counts on 200 MW. The missing 50 MW is about a quarter of the expected revenue, and the loan was sized on that revenue.
Two examples from RAI’s demo project. Each document is consistent on its own page; the problem only shows up when you compare them.
Cross-checking a dossier is careful work done by hand, against a deadline, and the hardest problems to catch are the documents that aren’t there at all.
Each document is written by a different party. Contradictions only surface when someone holds every figure in hundreds of pages against every other one.
Under the 2025 federal tax law, a solar project that starts construction after July 4, 2026 must be in service by the end of 2027 to keep its federal tax credit. A slow review eats into that window.
A thin dossier looks just like a full one. No irradiance study, no proof of site control: gaps only show up if someone checks against what should be there.
RAI does that cross-checking in one run, cites every source, and flags what’s missing.
Instead of asking one big AI model to read everything, the team split the review into small specialist agents. Joseph and Kiran built that pipeline. My job was the other side of it: making nine agents’ work readable, checkable and actionable for one analyst.
Each agent gets one narrow, focused prompt instead of one model juggling legal, finance, grid and permitting at once.
Document readers and researchers run at the same time, one per document or topic, so a review finishes in one run.
Every agent returns a structured, typed result, so the next agent and the dashboard never have to guess what free text meant.
The red flags come from contradictions between documents, so one agent’s only job is to hold every number against every other one.
An analyst drops in a project’s documents and gets back a scored, cited report. Here’s what happens in between, and where the person stays in charge.

Drop in the files, name the project, pick Fast or Deep diligence.

Each agent gets its own status box and narrates as it goes, so a long run never looks frozen.

Mid-run, the pipeline pauses on the gaps it found. The analyst picks which ones the agents chase. If no one answers before the timer runs out, every gap gets chased.
Read the score, dig into any finding, then ask questions or export the memo.



Clips recorded from the project repo running locally on its built-in demo data, with the agent backend offline.
Every finding links back to the document page or public source it came from. Nothing is asserted without a citation you can open.
On-track items stay grey. A screen full of green checkmarks teaches people to stop reading, so only what needs attention gets color.
If a source didn’t load, the UI says “fetch failed” or “not verified yet.” It never invents a link or shows an empty result as a real one.
Principle 03 in the actual screens. An analyst can’t act on a finding they can’t check, so every uncertain state is visible instead of smoothed over.

A claim links out only when the pipeline returned a real URL. Anything else carries an “unverified — no external source” tag. Statute citations link only to official pages that were checked to resolve.

Fields the report never surfaced read “Not in the report.” RAI doesn’t fill them with a plausible guess.

Live answers the report doesn’t support are marked “Not covered by this project’s report.” With no run behind a project, the rail says it can only answer from the findings.

If a live run fails, the analyst is told no results were saved, then chooses to retry or continue on demo data. Demo data only ever appears under a “Backend offline” banner.
Screens captured from the project repo in demo mode. The failed run uses the app’s built-in error simulation.
I designed the product and built the entire frontend: every screen, from uploading documents to the shareable memo.
The frontend is built against one typed report contract, so live results drop in unchanged. While a run is going, each agent narrates in its own status box.
RAI started at a hackathon, and we kept building it for about two weeks after. It was my deepest look yet at agentic architecture, and at what the engineers on our team were actually building.
While the agents ran, a prompt asked whether to keep waiting or cancel. It kept popping up on runs that were still working, and without solid checks in the code, it sometimes triggered the cancel.
What was actually taking too long? The prompt wasn’t reading the one signal that mattered: the live stream of agent activity on the page.
Tie the timer to that stream. Every new line of agent activity restarts a 3-minute timer, so the prompt only appears after real silence. The planned pause for gap review never counts.
The page’s behavior has to reflect what the backend is really doing. Once you’re past the mockup, a waiting screen is a reading of the live system.
Once I understood how the agents were built and how they worked through the documents, I could hand the engineers designs that fit what the system could actually do.
Designing for people working with agents that parse documents means designing the waiting, the sources and the uncertainty, not only the final answer.
Still to do: the human review bar is designed and built, but not wired into the project view yet. Next is mounting review, then saving every approved correction so reviewers make the agents better over time.