Mohammad Al Abdullah

Portfolio

Data product · 2026

Multi-source lead engine with CRM suppression

Design and build

A lead pipeline for a client that finds candidates across four sources, confirms every field against at least two of them, and refuses to hand over anyone already sitting in the client's CRM.

4

sources cross-referenced per candidate

57k

CRM contact records checked against every run

2

sources must agree before a field is confirmed

0

records written to the CRM without human approval

The brief

The client's problem was not finding leads, it was trusting them. Bought contact data arrives confident and partly wrong, and a sales team that gets burned twice stops using the list. On top of that, contacting someone who is already a client, or already said no, costs more than never contacting them at all.

The engine discovers candidates by position, company and location, then runs each one through a cascade where every source that follows both verifies what is already there and fills what is missing. A field counts as confirmed only when at least two sources agree.

It ends by cross-referencing the client's own CRM, which holds around 57,000 contact records, and either drops the match or shows it as a field-level difference for a human to approve.

Deliverables

What was handed over.

Cascade pipeline

Discovery, company-website reading, people enrichment, email verification and CRM suppression, each stage able to be switched off without breaking the ones after it.

Field-level provenance and confidence scoring

Every field on every candidate carries the trail of which sources produced it, and a confirmed badge counting only the sources that actually ran.

CRM suppression and approval flow

A weekly full mirror of the client's CRM, an indexed match on email then name plus company then name alone, and a decision on each match that a human signs off before anything is written.

Web dashboard and scheduled runs

A single-page dashboard for searching, inspecting the source trail behind any candidate and exporting, with searches dispatched as scheduled jobs.

Self-evaluation harness

A scoring pass that replays the website-reading stage over past runs and reports how often a delivered lead is actually named on their own employer's site, using no paid calls.

Actions

What I did.

  • Designed the cascade so a switched-off source is invisible rather than broken: never called, never warned about, and absent from both sides of the confirmation ratio, so turning one back on widens the ratio instead of retroactively condemning every lead gathered while it was off.
  • Put the free source before the paid one. The engine reads a company's own public contact pages first, and any person whose email is found there is dropped from the list handed to the paid enrichment source, so the reveal is never bought.
  • Made match strength gate the action, not just the decision: a weak match may drop a lead but is never allowed to write to an existing record.
  • Built the approval gate as two independent locks, both of which must be open before anything reaches the client's CRM.
  • Wrote the evaluation harness so the question of whether it works is answered with a number rather than an impression.

Problems and solutions

What went wrong, and what was done about it.

Every project has these. They are more informative than the finished result, so they are on the page.

Problem

A month's budget of paid enrichment credits was emptied in a single day.

Solution

Two guards, both asserted in the test suite so they cannot quietly regress. The paid source was restricted to enriching people already found rather than discovering new ones, which means no names in equals no call out and no spend, and a per-company cap was set to match the one lead per company actually delivered. Letting it discover again is worth around 38 percent more leads at several times the burn, which is a business decision to be taken deliberately, not a default to drift into.

Problem

The CRM offers no search endpoint and no changed-since endpoint, so answering whether a person already exists meant paging the entire database on every run.

Solution

Moved the full pull to a scheduled weekly job that builds and caches an index, so weekday runs cross-reference from memory instead of the network. An empty index is treated as a hard error rather than as a clean result, because the failure mode of a silently empty suppression list is contacting every existing client at once.

Problem

Real contact data was committed to git history once, putting a small number of real email addresses and phone numbers into a permanent record.

Solution

The data directories are gitignored, the database is the store of record rather than any file, and no scheduled job commits anything except a push receipt carrying counts and a timestamp and no personal data at all. Which individuals were delivered is answerable inside the client's own CRM, which has a retention policy, rather than in git history, which does not. The historical exposure is documented in the open rather than quietly cleaned, because removing it from history is a separate decision that has not been made.

Problem

The database shares an instance with an unrelated live product, so a restore to undo a mistake here would roll that product back too.

Solution

Migrations are forward-only and additive with no drops and no renames, and the runner refuses any migration file that names the shared schema, with that refusal asserted in the test suite. The safety comes from the migration being unable to reach the other product, not from remembering to be careful.

Problem

The dashboard had reset and seed buttons sitting one mis-click away from wiping four tables of the client's live data.

Solution

Removed from the interface and replaced with command-line operations that each require an explicit target. There is no wipe-everything path left to click.

Outcome

A lead list the client's sales team can act on without checking it first, where every field carries its own evidence, nobody already in the CRM gets contacted twice, and nothing is written back without a human approving the specific rows.

What it took

TypeScriptNodePostgreSQLSupabaseREST API integrationScheduled jobsData reconciliation