Evaluating a data product by looking at the sample response is how people end up with a feed that fails badly in eight months. The sample is always the happy path. These questions are about everything else.
On freshness
When was this number measured, as opposed to when was the response built? Does every reading carry its age? And is there a difference in the response between something you computed and something you passed through from a third party?
On failure
What does an outage look like from my side — an error, or the last good value, or an empty result? Are the error codes stable and documented? Is the rate limit counted per key or per address?
On scope
What is in the universe being scanned, and what is filtered out of it? For any percentile or ranking, what window is it measured over? On a tiered plan, how do I tell a withheld field from a zero?
On history
How far back does the recorded series go, and can I query it? This is the one where vendors most often have a worse answer than their marketing implies, because the live computation is easy and the archive is not.
How to read the answers
Speed matters as much as content. A vendor who answers the freshness question immediately has thought about it; one who has to go and check has not, which means nobody has been maintaining the guarantee.
Any answer along the lines of it never really goes down is a non-answer. Everything goes down. The question was what it looks like when it does, and a vendor who has not designed for that has designed for it badly by default.