Mohammad Al Abdullah

Portfolio

Internal tooling · 2026

Multi-agent validation harness for feasibility work

Design and build

An adversarial review layer that audits a feasibility study before a client sees it, built because checking your own arithmetic is the one review that never finds anything.

6

specialist review domains

2

independent validators, logic and evidence

1

final auditor, no output ships without it

The brief

A feasibility study that reaches a client with an error in it costs more than the study was worth. The problem is structural: the person best placed to find the mistake is the person who made it, and they have already read the document too many times to see it.

The harness runs a study through specialist reviewers that do not share the author's assumptions, then through validators that check the reasoning and the evidence separately, and finally through an auditor that has to sign off. Its memory is plain text so every lesson it learns stays human-readable and reviewable.

Deliverables

What was handed over.

Domain review agents

Financial, market, technical, sustainability, risk and funding reviewers, each looking at the study through one lens rather than all of them at once.

Independent validators

Logic and evidence checked separately, because an argument can be internally consistent and still unsupported, and the two failures need different questions.

Reflection and durable memory

Every review writes back what it learned, in plain markdown, so a mistake caught once becomes a standing check rather than a lesson that leaves with the reviewer.

Final audit gate

Nothing is treated as finished until an auditor with the whole picture signs it off.

Actions

What I did.

  • Ran the haulage feasibility study through the harness as a validation pass, on top of the forensic audit it had already had.
  • Kept the memory layer as plain readable files rather than an opaque store, so a claim the system makes can be traced to the reflection that produced it.
  • Scoped a version of it specifically to fleet electrification, carrying the pilot facts, the maintenance rebase and the data-source rules.

Problems and solutions

What went wrong, and what was done about it.

Every project has these. They are more informative than the finished result, so they are on the page.

Problem

A reviewer given the whole document and asked to find problems returns a summary, not an audit.

Solution

Split the review by domain and gave each reviewer one question. A financial reviewer asked only about the financial case pushes on the discount rate; the same reviewer asked about everything writes an executive summary and misses it.

Problem

Lessons from a review evaporated as soon as it finished.

Solution

Reflections are written to durable, human-readable memory that later reviews retrieve before they start, so the second study inherits what the first one learned.

Outcome

The two pilot studies went to clients having been through a review that did not share their author's assumptions. The haulage study in particular carried seven versions, a forensic audit and an independent validation pass behind it, which is the reason its numbers survived the room they were presented in.

What it took

PythonMulti-agent orchestrationLLM workflowsVector retrievalEvidence validation