You type a question into a chat window. Somewhere between that and the answer, several discrete things happen, and every one of them is inspectable.

First, the catalogue

When the client connects, it asks the server for a list of tools. Each comes back with a name, a human-readable description, a schema describing its arguments and a set of hints about what it does. The model never sees your credentials or the server internals — just this catalogue.

If the server offers a free tier and a paid one, the catalogue usually lists everything, with the paid ones marked. The tools exist either way; what changes is what happens when they are called.

Then, the choice

The model reads your question and the catalogue and decides. It may call nothing, one tool, or several in sequence, using the result of one to shape the arguments of the next. This is why a question like whether a coin looks stretched can trigger three calls: one for the map, one for funding, one for the volatility percentile.

Then, the payload

The tool returns structured data, which is inserted into the conversation as text the model can read. This is the moment where honesty in the feed pays off. A response carrying freshness fields and named refusals gives the model something to say when data is missing. A response that quietly returns zeroes gives it nothing to notice.

Finally, the prose

The model writes an answer from the payloads and whatever else is in the conversation. A well-behaved one cites which tool produced which figure, so you can check it.

The failure to watch for is an answer that contains numbers no tool returned. That is the model filling a gap from its training data, and it is exactly as reliable as a half-remembered statistic.