Locus product guide · Updated
How to Evaluate AI Agent Tools and API Providers
Evaluate an agent tool by whether it returns useful data for a representative task, respects the expected request and response contract, and records the correct charge. Compare coverage, latency, price basis, access requirements, and recovery behavior. A catalog listing, successful connection, or HTTP success code alone does not establish task quality.
Read the listing’s coverage and provenance
The Locus catalog distinguishes reviewed providers, marketplace listings, and live-quoted external services. A total catalog count combines different kinds of access; it should not be read as a claim that every marketplace entry has the same review status or service guarantees.
Inspect the specific operation, input schema, data source, geographic or time coverage, and pricing basis. Some services need valid identifiers; others return a subset of results or an asynchronous job. A recognizable provider name is not a substitute for that operation-level check.
Define success before running the request
For company enrichment, require the intended company and the fields your workflow needs. For extraction, compare the returned content against a known page. For code execution, check stdout, stderr, exit status, and resource limits. Use non-sensitive test inputs you are authorized to send to the provider.
Include at least one known-positive request and one request that should return no match or an explicit validation error. This separates a working integration from one that returns empty output for every input.
- Result relevance and required fields
- Supported filters, geography, and freshness
- Latency measured through the intended access path
- Actual charge and unit basis
- Behavior for empty results, errors, and retries
Measure the task cost, not only a sample call
A tool may bill by request, token, page, generation, returned record, or live quote. One task can require multiple calls and follow-up operations. Record the final receipt for each step, including any release or retained charge, then total the workflow.
Price examples in documentation are explanations, not standing quotes. Read current effective pricing through the catalog and authorization flow. For live-quoted operations, use a supported maximum charge ceiling before dispatch.
Separate integration tests from provider acceptance
Sandbox tests can verify your parsing, attribution, and handling of billing states without spending credits or calling an upstream provider. They do not prove that the external service currently has the data or behavior you need. Validate the real paid path before customer rollout.
Keep a dated record of the environment, operation, expected result, observed result, and recorded charge. Repeat the relevant acceptance checks when your inputs, the provider contract, the operation schema, or the pricing basis changes. Publish benchmarks only when the methodology and evidence support the comparison.
Frequently asked questions
Does a verified provider label cover every marketplace tool?
No. Reviewed provider listings and marketplace services are distinct catalog tiers. Inspect the individual listing and operation rather than applying one label to the total catalog.
What is a useful first acceptance test?
Make a small authorized request with a known-positive input through the intended interface, assert the required useful result, and inspect the actual recorded charge and attribution.
Implementation references
Use these first-party references for current request contracts and account requirements. Tool availability, prices, and negotiated terms can change.