What Happens When Two AI Agents Disagree?

Disagreement is not a bug in a multi-agent system. It is a design question you answer before the conflict happens.

Two AI agents look at the same refund request.

The support agent sees a customer who has been with the company for five years and has spent fourteen thousand dollars. It recommends approving the one hundred and fifty dollar refund.

The finance agent reads the refund policy, sees a limit of one hundred dollars per incident, and recommends rejecting it.

Neither agent is broken. Neither one ignored its instructions. They optimized for different things and arrived at opposite answers.

So what happens next?

https://youtu.be/Vr-teESQekQ

That question is the real work in multi-agent systems. Not building more agents. Deciding what happens when the agents you already have do not agree.

Disagreement usually means multiple valid goals

A single agent is easy to reason about. You give it one objective and it pushes in one direction.

A business does not have one objective. It has several, and they pull against each other all day:

Customer experience wants the refund approved. Cost control wants it refused. Risk wants more evidence. Compliance wants the written policy followed. Revenue wants the customer retained. Speed wants an answer in two seconds.

Humans resolve this constantly, mostly without noticing. A support lead weighs loyalty against policy and makes a call. If the call is unusual, she walks over to her manager.

When you turn those roles into agents, the weighing does not disappear. It just becomes invisible unless you design for it.

Here is a useful test. If your agents never disagree, you probably do not have a multi-agent system. You have one agent and some assistants that approve whatever it says.

Before you resolve it, check whether it is really a disagreement

This is the step most teams skip, and it is the one data engineers should care about most.

Ask a simple question first. Are both agents looking at the same facts?

Imagine the finance agent read the refund policy document that was current last month, while the support agent read the version updated last week. The limit changed from one hundred to two hundred dollars. The two agents now give different answers, and it looks like a judgment conflict.

It is not. It is a context problem.

The same thing happens with definitions. The support agent calculates lifetime value from raw order rows, including cancelled orders. The finance agent uses the approved customer value table, which excludes them. One sees fourteen thousand dollars. The other sees nine thousand.

And it happens with freshness. One agent reads a table that was refreshed an hour ago. The other reads a copy that stopped updating two days ago when a pipeline failed quietly.

So the first rule of multi-agent design is not about agents at all:

Resolve context before you resolve conflict. If two agents are reading different versions of the truth, no arbitration logic will save you. It will just pick a winner, confidently, from bad inputs.

This is why a governed catalog, one official definition per metric, and reliable freshness checks are not paperwork. They are the thing that makes agent disagreement meaningful rather than random.

Authority has to be decided before the conflict

Once you are sure both agents see the same facts, you still need a way to decide.

The wrong approach is to let the agents argue until one of them sounds more convincing. Confidence is not authority. A well written paragraph is not a decision right.

The better approach is boring and written down in advance. Something like:

Compliance overrides pricing. Security overrides convenience. A hard policy limit overrides a retention recommendation. A fraud signal can block a transaction that every other agent wants to approve.

Notice what that list really is. It is not AI design. It is the same delegation of authority a company already writes for its people, expressed in a form a system can apply.

If nobody can tell you who wins when the refund agent and the policy agent disagree, the system does not have a governance gap in the future. It has one now.

Escalation is a design choice, not a failure

The second mistake is treating "the agents could not agree" as an error state.

In a well designed system, escalation is a normal path, and it has steps:

Compare the evidence. Are both sides citing the same records? Is one using a stale source?

Check confidence. If one agent is sure and the other is guessing, that matters. If both are unsure, that matters more.

Apply the written policy. Deterministic rules beat persuasion. A hard limit is a hard limit.

Hand it to a human. If the conflict cannot be resolved safely, stop and ask. That is not the system failing. That is the system working exactly as intended.

The point of escalation is to make sure that the cases nobody designed for end up in front of someone who can think, rather than being quietly resolved by whichever agent spoke last.

A simple risk pattern that works

You do not need an elaborate framework to start. Sorting decisions by risk gets you most of the way:

Low risk resolves automatically. A twelve dollar refund on a clear duplicate charge. Let the system settle it and log the decision.

Medium risk gets a validation step. A one hundred and fifty dollar refund outside policy. The system must re-check the facts, confirm the policy version, and record why it chose what it chose.

High risk goes to a human. A three thousand dollar credit, a contract exception, anything touching compliance or a regulated customer. The system prepares the case, gathers the evidence, and waits.

Most teams get this backwards. They automate the hard cases because those are the interesting ones, and leave the easy volume to people. Do the opposite. Automate the boring majority and buy your people time for the cases that actually need judgment.

Why shared state matters more than clever coordination

Multi-agent systems usually fail in coordination, not reasoning.

Two agents work on the same case at the same time and neither knows what the other did. One issues a credit while the other opens a dispute. The customer gets two emails that contradict each other.

The fix is not a smarter prompt. It is shared state that every agent reads and writes:

What case is this? What has already been done? Who is currently working on it? What was decided, by whom, and on what evidence?

That is a table. A well designed, well governed, auditable table. Which is a quietly reassuring thing for data engineers to hear, because it means the hardest part of the agentic era is work you already know how to do.

Try it in Databricks Free Edition

Here is a small exercise you can run in Free Edition in about fifteen minutes. It builds the shared decision log that makes disagreement visible instead of invisible.

Create a table where every agent records its recommendation.

from pyspark.sql import functions as F

agent_recommendations = spark.createDataFrame(
    [
        ("CASE-1001", "support_agent",  "APPROVE", 0.88, "loyal customer, 5 years"),
        ("CASE-1001", "finance_agent",  "REJECT",  0.95, "exceeds 100 usd policy limit"),
        ("CASE-1002", "support_agent",  "APPROVE", 0.91, "duplicate charge confirmed"),
        ("CASE-1002", "finance_agent",  "APPROVE", 0.93, "within policy limit"),
    ],
    ["case_id", "agent_name", "recommendation", "confidence", "reason"],
)

agent_recommendations.write.mode("overwrite").saveAsTable("agent_recommendations")

Now find the cases where the agents disagree.

SELECT
  case_id,
  COUNT(DISTINCT recommendation) AS distinct_answers,
  COLLECT_SET(recommendation) AS answers
FROM agent_recommendations
GROUP BY case_id
HAVING COUNT(DISTINCT recommendation) > 1

CASE-1002 is agreed and can settle automatically. CASE-1001 is contested.

Now apply authority instead of confidence. The policy agent holds the veto on policy limits, so its rejection stands, and the case is marked for review.

SELECT
  case_id,
  CASE
    WHEN MAX(CASE WHEN agent_name = 'finance_agent'
                   AND recommendation = 'REJECT' THEN 1 ELSE 0 END) = 1
      THEN 'BLOCKED_BY_POLICY'
    WHEN COUNT(DISTINCT recommendation) = 1
      THEN 'AUTO_RESOLVED'
    ELSE 'NEEDS_HUMAN_REVIEW'
  END AS outcome
FROM agent_recommendations
GROUP BY case_id

Two things are worth noticing.

The deciding logic is deterministic. It does not depend on which agent wrote a better sentence.

And every decision now has a record: who recommended what, how sure they were, why, and what the system did about it. Run DESCRIBE HISTORY agent_recommendations and you can see when each version was written.

That record is the difference between a collection of agents and a governed agent system.

The design questions that actually matter

When someone describes their multi-agent plan, these four questions tell you how far along they really are:

Who has authority? When two agents disagree, which one wins, and is that written down anywhere?

Which source is trusted? If two agents report different numbers for the same thing, which table is the official one?

When should the system stop? What value, what risk level, what category means no autonomous action?

When must a human step in? Not "can a human intervene", but at which points is human approval required before anything happens.

If the answers are clear, the system can be trusted with real decisions. If they are vague, the system will still make decisions. You just will not know how.

Where this fits the Context Advantage

This is Context and Control working together.

Context is the shared ground truth: the same facts, the same definitions, the same current state, available to every agent that needs them.

Control is the structure on top: who has authority, what requires approval, where the system stops, what gets recorded.

Intelligence is becoming distributed. Many agents, many models, many tools, many decisions happening in parallel. But accountability does not distribute. It still has to live somewhere specific, with a name attached.

The goal was never to build agents that always agree. Agents that always agree are just an echo. The goal is a system where disagreement surfaces early, resolves against written rules, and reaches a person when it should.

Keep learning, keep building, keep growing.

Brick by brick.