Every model call is metered and aggregated from call to step to scenario to suite, so each run answers both halves of the question.
Autonomous QA is only worth adopting if it is cheaper than the alternative, and that is an empirical question most tools cannot answer about themselves. Argus treats cost as a first-class output: every model call is recorded and rolled up, and a cheap model runs everything first with only agent-side failures escalated to a stronger one — each tier costed separately, so you can see what escalation bought.
Every model call is metered, including router and verifier calls.
Cost aggregates from call to step to scenario to suite to release.
A cheap model runs first; only agent-side failures escalate.
Each tier is costed independently.
Optimises cost per successful test, not just pass rate.
Only failures with an agent-side cause — never product or assertion failures.
It is the number that decides whether autonomous QA is actually cheaper than the alternative.
Yes. Grounding space, reasoning effort and provider routing are configuration, not code.
A scenario is an objective, a role and assertions — no selectors, nothing to rewrite after a redesign.
Creator, approver and supplier each get a fresh, independently authenticated session.
A lost agent is never filed as a product defect.
A passing run distils its path into intents, never coordinates, for later runs to reuse.
See the complete product: problem, features, how it works and deployment.
Request a demo to watch a multi-role regression suite run through your real interface.