A rate limit is a cap on how many requests a server will answer from one caller in a stretch of time. Go over it and the server refuses until the window resets. People building scripts know this well. People using agents often meet it for the first time when the agent stalls halfway through an answer.
Why agents hit limits
A script calls what you wrote. An agent calls what it decides to, and it decides again on every turn. Ask a follow-up and it may fetch the same levels a second time. Ask it to compare three tickers and it may make a call per ticker when one would do. It is being thorough, and it has no sense of cost unless you give it one.
What a refusal looks like
A well-built server refuses in a way a program can read. The standard signal is a too-many-requests status with a time to wait. A stopped feed on NoVo’s server answers with an error and a retry time. The agent’s job is to respect it. Tell it to wait, or to report that the tool is resting. It should never retry in a loop. The design of the ceilings themselves is covered in rate ceilings are a product decision.
Do not ask faster than the data changes
Every reading carries age_seconds. If the agent fetched the levels a minute ago, a second call will mostly return the same thing. Tell the agent to reuse a reading it already holds in the conversation unless you ask for a refresh. When you do ask, it can compare the new as_of with the old one and tell you whether anything changed.
The delayed free tools make this plainer still. A reading that is delayed by design does not get fresher because you asked twice.
Prefer one call that covers the question
Some tools are built to answer a whole question at once. get_ticker_brief gathers everything the free surface holds on one name in a single call. get_stocks_onchain with no symbol returns the largest perps and tokens together. get_dealer_levels covers SPY, QQQ and IWM. An agent that knows this spends one call where it might have spent five.
Know the separate allowance
On the paid plans ask_novo is counted on its own daily allowance, apart from the data tools. A read from Dr. NoVo, the Financial Markets Super Intelligence, takes real work to produce. So an agent should not route every small lookup through him. Save those calls for questions that need interpretation, as in letting Dr. NoVo answer or calling the data tools.
Scheduled agents and standing rules
An agent that runs on a timer needs the most care. Work out its calls per run and its runs per day before you switch it on. Add room for retries, because a bad hour is when retries pile up. Then set the timer no faster than the data moves. A loop that polls a five-minute reading every few seconds spends budget and learns nothing.
Three lines cover it. Reuse readings already in this conversation. Use the widest single tool that answers the question. If a tool gives a retry time, wait and tell me. The plans are listed on the MCP & API page, and what a key does and does not open is in your API key says who you are.