Cost accounting

What did this suite cost, and what did it find?

Every model call is metered and aggregated from call to step to scenario to suite, so each run answers both halves of the question.

Overview

Cost per test

Autonomous QA is only worth adopting if it is cheaper than the alternative, and that is an empirical question most tools cannot answer about themselves. Argus treats cost as a first-class output: every model call is recorded and rolled up, and a cheap model runs everything first with only agent-side failures escalated to a stronger one — each tier costed separately, so you can see what escalation bought.

What it does

The economics, measured

Every model call is metered, including router and verifier calls.

Cost aggregates from call to step to scenario to suite to release.

A cheap model runs first; only agent-side failures escalate.

Each tier is costed independently.

Optimises cost per successful test, not just pass rate.

How it works

How cost is tracked

1Each model call records its tokens and price
2Costs roll up per step and per scenario
3Escalations to the stronger model are costed separately
4The suite report states what the run cost
FAQ

Common questions

What gets escalated to the stronger model?

Only failures with an agent-side cause — never product or assertion failures.

Why measure cost per successful test?

It is the number that decides whether autonomous QA is actually cheaper than the alternative.

Can we change the model?

Yes. Grounding space, reasoning effort and provider routing are configuration, not code.

Explore more

More of what Argus QA does

Get started

See agents test your own application

Request a demo to watch a multi-role regression suite run through your real interface.