One thing is measured here: whether the assistant picks the right query and the right arguments for the question asked. It is not a score for the factual accuracy of every answer — the figures come from the data itself, and this page checks whether the question reached it.
Back to the chat