Peregryn™ · Identity Resolution Engine

Match 12 million records in under seven minutes — and defend every merge.

Peregryn resolves identities across B2C customers, suppliers, and B2B accounts. Deterministic rules and probabilistic scoring — Fellegi-Sunter class — run on a DAMA-DMBOK-aligned pipeline: profile, cleanse, standardize, match, merge. Every merge is reversible at the cluster level. It runs on commodity x86, entirely on your premises. No GPU. No cloud.

The problem

Every system swears the customer is unique. The merge history says otherwise.

Duplicate parties accumulate quietly — a customer keyed three ways in three systems, a supplier that is also an account, an acquisition that doubled the estate overnight. When the forecast misses, the exam letter arrives, or the integration stalls, the question is no longer whether duplicates exist. It is whether you can resolve them — and reconstruct every decision that resolution made.

What it does

One resolution engine across customers, suppliers, and accounts.

01 · Match

Two disciplines, one decision

Deterministic rules decide what must match; Fellegi-Sunter-class probabilistic scoring decides what probably does. A match is a scored decision with stated evidence, not a hunch.

02 · Survive

Survivorship you can read

When records merge, explicit survivorship rules decide which values persist on the surviving record — not load order, not luck. The golden record is an argument you can reconstruct.

03 · Reverse

A merge that can be undone

Merge and unmerge are first-class operations, with full rollback at the cluster level. An integration decision is never a one-way door.

B2C customers, suppliers, and B2B accounts flow through one governed pipeline — profile, cleanse, standardize, match, merge — aligned to DAMA-DMBOK. One pipeline, five stages, no side doors.

Performance

Built to resolve the whole estate, not a sample.

Measured on commodity x86 hardware — no accelerator, no cluster, no external service in the path.

12 million records, matchedprofile → merge · end-to-end · measured
<7 minutes
GPUs requiredcommodity x86 only
0 GPUs
Cloud services in the pipelineprofile through merge
0 services
Deploymentruns where your data lives
100% on-prem

Stated plainly

The twelve-million-record figure is a measured end-to-end run on commodity x86 hardware. We will reproduce the run on your own data during a proof-of-value engagement — a benchmark you cannot reproduce is marketing, and this one is not.

Control

No merge is a one-way door.

  • DefaultMatch candidates are scored and recorded first. Nothing merges until the thresholds you set say it should.
  • Cluster rollbackAny merged cluster can be rolled back in full. Unmerge restores the pre-merge state — not an approximation of it.
  • Golden recordExplicit rules decide which values survive a merge. The rules are reviewable by your governance function before anything runs.
  • ScopeOne engine across B2C customers, suppliers, and B2B accounts — the three places the same real-world party hides under different keys.

PIPELINE · PROFILE → CLEANSE → STANDARDIZE → MATCH → MERGE · DAMA-DMBOK-ALIGNED

Why Pacific Data Integrators

Fifteen years inside the systems your data lives in.

Peregryn comes from a firm with 15+ years and 100+ implementations across Informatica, Salesforce, Snowflake, and Databricks — for banks, insurers, healthcare, government, and retail. It is the same engine discipline behind ForecastGuard for Salesforce: match decisions a CRO can defend to a board and a CDO can defend to an examiner.