ApisTrust
Independent quality measurement for the x402 agent economy
Last measured 2026-08-12

Methodology

A rating is only worth what its method is worth. This page describes exactly how the numbers on this site come about, including what we throw away and what we cannot see.

Source

The full catalogue is read every night from the public x402 discovery API (api.cdp.coinbase.com). No sampling: every listed endpoint is included. Additions, removals, price changes and changes of receiving address are recorded.

Measurement

Each endpoint receives one GET request, and a POST retry where the endpoint expects one. Timeout is eight seconds. We send the user agent ApisTrust/1.0; we do not disguise our traffic. No payment is made during this stage, so the measurement costs providers nothing beyond one request.

An endpoint counts as protocol-compliant when it answers HTTP 402 with machine-readable payment details. Both protocol versions are accepted: version 1 carries them in the JSON body, version 2 in the payment-required header. We record the price actually demanded and compare it with the catalogue entry.

Scoring

The ApisTrust Score is the share of valid runs in which an endpoint answered correctly, on a scale of 0 to 100. A penalty of 15 points applies when an endpoint demands more than its listed price, and 30 points when it demands more than double. Endpoints with fewer than five measurements are not scored. A provider's score is the average across its endpoints.

An outage counts only when two consecutive valid runs fail. A single miss does not lower a score, because a single miss is as likely to be our fault as theirs.

Discarded runs

Entire runs are excluded when their share of pure network errors — timeouts, DNS failures, refused connections, TLS errors — exceeds five times the normal level and lies more than one percentage point above it. In practice that level sits between 0.06 and 0.3 percent; on 10 August 2026 it reached 6.5 percent, and that run was discarded. Without this rule 827 providers would have been marked unreliable for a fault that was ours.

Excluded runs are always named on the overview page, with their figures.

Paid sampling

Availability says nothing about whether an answer is correct. Once a week we buy answers from selected providers using test cases whose correct outcome we already know. A service that warns about malicious tokens is given a token proven malicious by two independent sources; a sanction screening service is given an address on the OFAC list. A single source is never treated as ground truth, because single sources produce false positives.

Known limits

We measure from one location in Germany. A provider unreachable from there may be reachable elsewhere. We measure once per day, so short outages between runs are invisible. Paid sampling covers a small selection, not the whole market, and providers whose price exceeds our per-request budget are not sampled at all. Endpoints appearing in the catalogue for the first time carry no score until five measurements exist.

Corrections

Every figure here is reproducible with the same public sources. Providers who believe a measurement is wrong are invited to write to contact@apistrust.com; corrections are published alongside the original data rather than replacing it silently.