If you use more than one AI app, you will sooner or later ask both the same market question. The answers will not match. It is tempting to trust the one that reads better. The better test is to find out why they differ, and the cause is usually one of five.

They read at different times

A reading has a time. Two calls a few minutes apart can return different values, and both are correct. Check the as_of on each answer. If neither agent gave one, ask. Until you have the two times, you cannot say the answers disagree at all.

They called different tools

On NoVo’s server, the same subject can be reached through more than one tool. get_dealer_levels serves the dealer map delayed and with some fields gated. get_equity_map, a paid tool, serves it whole and current. An agent with a key and an agent without one will return different walls on a fast day. Ask each agent which tool it used.

The same applies to index rows. One agent may quote SPY from the dealer levels. Another may quote a row from the quotes tool that is a future. Those are different instruments at different prices.

One of them did not call a tool

This is the cause that matters most. One agent fetched the figure. The other wrote from memory and sounded just as sure. Ask each to name the tool and the age. The one that cannot is the one to set aside. The pattern is described in the number the model made up.

One dropped or blurred something

Both called the same tool at the same time. One says the gamma flip is withheld on the free tier. The other says there is no flip today. The payload was identical. One agent read the gated field and the other did not. The first answer is right.

One agent quotes the wall at the strike. The other says near a round number. One lists funding per venue. The other gives a single blended figure that no venue shows. Rounding is harmless once you see it. Blending is a fault, and the agent that kept the venues apart has the better answer.

How to settle it

Go to the payload. Most AI apps will show the tool call and what came back. Read the field yourself. The answer that matches the payload wins, however plainly it was written. Smooth prose is not evidence, a case made at length in fluency is not evidence.

If you cannot see the payload, ask both agents the same follow-up: quote the exact field and its value. An agent that fetched the data can do this. An agent that recalled it cannot.

When they agree

Agreement is weaker proof than it looks. Two models can share the same stale memory and repeat it in unison. Two answers that match and name no tool are still two unsourced answers. Two answers that match, name the same tool and give close times are worth trusting.

Comparing agents is a cheap audit. It costs a second question. Done now and then, it shows you which of your setups carries the time, the source and the caveats through. Both agents can read the same server at the MCP & API, so the payload is common ground. The habit behind all of this is in an answer you cannot check is an opinion.