Skip to main content
← Back to BlogIdentity Resolution Is Quietly Becoming an AI Problem, Not Just a Data Problem

Identity Resolution Is Quietly Becoming an AI Problem, Not Just a Data Problem

AIHelpTools TeamSeptember 25, 2026
identity-resolutioncustomer-data-platformllmdata-engineeringmartech

Table of Contents

  1. Why Identity Resolution Suddenly Matters Again
  2. The Old Way: Batch Jobs and Business Rules
  3. Why LLMs Change the Game
  4. The New Stack: AI-First Identity Resolution
  5. Before and After: A Real Architecture Comparison
  6. The Synthetic Traffic Problem Nobody's Talking About
  7. What This Means for Your Data Platform

Why Identity Resolution Suddenly Matters Again

Identity resolution was supposed to be a solved problem. CDPs conquered it between 2016 and 2020. Match emails, stitch device IDs, build a golden record. Done.

Except now we have AI agents making real-time decisions about what to show customers. An agent can't wait for your nightly batch job to finish. It needs a profile lookup in under a second. And that profile needs to be correct, not just consistent.

The gap between what your identity graph can deliver and what your AI systems need is wider than most teams realize. You built your identity resolution for reporting dashboards and email campaigns. You didn't build it for an LLM that's orchestrating a personalized video sequence at runtime.

Analogy: Your identity resolution system is like a library card catalog. It works great when patrons visit once a day and fill out request slips. It completely falls apart when a thousand people need book recommendations simultaneously, and half of them are asking in languages the catalog was never designed to handle.

The Old Way: Batch Jobs and Business Rules

Traditional identity resolution runs on deterministic rules and scheduled jobs. You write SQL that matches on email, then phone, then device ID. You run it nightly, maybe hourly if you're fancy. You handle edge cases with business logic: if the last name changed, check for marriage records. If the email domain shifted, maybe it's a job change.

This approach scales poorly in two directions. First, it can't handle the volume of lookups an AI agent generates. Second, it can't make fuzzy decisions about ambiguous matches without someone writing explicit code for every scenario.

Here's what a typical 2023 identity resolution pipeline looked like:

  • Raw events land in a data lake
  • Nightly Spark job runs matching rules
  • Results written to a customer 360 table
  • Applications query that table (with 12-24 hour lag)
  • Edge cases escalated to data ops team

The bottleneck isn't the compute. It's the decision-making framework. Every ambiguous case requires human judgment encoded as code. Every new data source requires new matching rules. Every market with different naming conventions needs custom logic.

Why LLMs Change the Game

LLMs are genuinely good at fuzzy matching and disambiguation. Not because they're magic, but because they've seen billions of examples of how names, addresses, and identifiers vary in the real world.

An LLM can look at "Robert Smith, rob.smith@email.com, 555-1234" and "Bob Smith, rsmith@email.com, 555-1234" and make a confident probabilistic judgment. It doesn't need you to write a rule that says "Robert can be shortened to Rob." It already knows.

More importantly, LLMs can explain their reasoning. When a match is ambiguous, a traditional rules engine just fails or makes a binary choice. An LLM can return: "85% confidence these are the same person based on name similarity, phone match, and email pattern. Email domain change suggests job transition."

That contextual explanation matters when you're feeding identity data into downstream AI agents. The agent can adjust its behavior based on match confidence.

The New Stack: AI-First Identity Resolution

The architecture is flipping. Instead of batch jobs feeding static tables, you're building a streaming system with an LLM in the loop.

Event Stream Vector Search LLM Match Layer Identity API

AI-first identity resolution stack

Here's what changes:

Event Stream: Customer actions flow in real-time, not batched overnight.

Vector Search: Profile attributes get embedded into vector space. Similar profiles cluster together naturally. This handles the "fuzzy match" problem at scale.

LLM Match Layer: When vector search returns candidates, the LLM makes the final call. It weighs evidence, handles ambiguity, and provides confidence scores.

Identity API: Sub-second lookups for downstream agents. The API returns not just "who is this" but "how confident are we" and "what's the reasoning."

Before and After: A Real Architecture Comparison

DimensionTraditional ApproachAI-First Approach
Match LogicHard-coded SQL rulesLLM inference on candidate pairs
Update FrequencyNightly or hourly batchReal-time streaming
Ambiguity HandlingEscalate to data opsConfidence scores with reasoning
New Data SourceWrite new matching rulesAdd to embedding model training
Query LatencyMinutes to hoursMilliseconds
Match Accuracy75-85% (deterministic)85-95% (probabilistic)
Explainability"Rule #47 triggered"Natural language reasoning
Cost ModelFixed compute schedulePay per inference

The Synthetic Traffic Problem Nobody's Talking About

Here's where it gets weird. AI agents don't just consume identity data. They generate it.

When an LLM-powered agent browses your site on behalf of a customer, is that the customer's session or the agent's session? When an AI assistant adds items to a cart, whose intent are you tracking?

Brian Silver from ZoomInfo calls this "synthetic traffic." Agents acting on behalf of humans create new identity records that need to be reconciled back to the actual person. Your identity resolution system now needs to:

  • Detect agent traffic versus human traffic
  • Link agent actions to the human principal
  • Decide when agent behavior should update the human profile
  • Handle cases where the agent is acting autonomously

None of the traditional identity resolution frameworks were designed for this. You need an AI-native approach that can reason about agency and attribution.

What This Means for Your Data Platform

If you're running a CDP or customer data platform, this shift has immediate implications:

Latency becomes critical. Your identity graph needs to support sub-second lookups, not hourly refreshes. That means moving from batch to streaming, probably rebuilding your storage layer.

Probabilistic beats deterministic. Stop trying to write perfect rules. Embrace confidence scores. Let downstream systems decide how to handle uncertain matches.

Explainability is non-negotiable. When an LLM makes a match decision, you need to know why. Not just for debugging, but for compliance and trust.

The cost model flips. You're trading scheduled compute for inference costs. Budget accordingly. A million identity lookups per day at $0.002 per call adds up fast.

PII sprawl gets worse before it gets better. LLMs need context to make good matches. That means more PII in more places. Your data governance team will not be happy.

The good news: you don't need to rebuild everything overnight. Start with a hybrid approach. Keep your batch jobs running. Add an LLM layer for ambiguous cases. Measure the improvement. Scale up as you prove value.

The Brutal Truth

Identity resolution was never truly solved. We just built systems that were good enough for the use cases we had in 2020. Those systems are not good enough for AI agents that need instant, confident answers about who someone is.

The irony is thick. We're using AI to solve the identity problems that AI systems created. LLMs need better identity resolution to work properly, and LLMs are the best tool we have for doing identity resolution at the speed and scale AI requires.

If your data platform roadmap doesn't include rethinking identity resolution for the AI era, you're optimizing for a world that's already gone. The batch jobs that powered your customer 360 in 2023 are not going to cut it when you have a dozen AI agents per customer, all needing real-time profile lookups.

Start with the hard problem. Not the AI features. Not the fancy agent orchestration. Fix identity resolution first. Everything else depends on it.