Glossary · Argus QA

What is test evidence?

What is test evidence? A clear explanation for Azerbaijani business — and how Argus QA applies it.

Comprehensive Test Evidence in AI Testing

Test evidence serves as the definitive documented proof that a software test was executed and whether it successfully met its defined objectives. In the context of autonomous regression testing for enterprise web applications, evidence evolves beyond simple pass/fail binary results. It encompasses a granular audit trail where AI computer-use agents record every observation, the underlying reasoning for each decision, and the specific action taken during every turn of the execution process. By leveraging a system where every step is persisted regardless of the outcome, organizations gain a transparent view into the agent's behavior. This approach transforms testing from a 'black box' process into a verifiable sequence of events. By combining screenshots with step-by-step reasoning and independent assertion verification, the platform ensures that the evidence gathered is objective, reproducible, and sufficient for high-stakes enterprise quality assurance.

Capabilities

Strategic Advantages of Detailed Evidence

Accelerates root-cause analysis by providing a complete record of agent reasoning and action logs for every step

Eliminates bias and false positives through the use of an independent verifier to check assertions

Prevents test interference in shared environments using run-scoped test data naming to avoid collisions

Provides granular financial visibility via metered model execution tracking from the call level up to the release suite

Ensures high-fidelity bug reporting by flagging suspected defects for human review rather than autonomous filing

Maintains strict environment traceability by recording the specific build stamp for every execution run

Argus QA Evidence Capture Capabilities

Persistent Action Logs

Every step is persisted regardless of outcome. The agent observes a screenshot, states its reasoning, and records one action per turn.

Independent Verification

To ensure objectivity, assertions are checked by an independent verifier rather than the agent that performed the work.

Failure Classification

Failures are categorized into specific types: product defect, agent failure, environment failure, assertion failure, loop, step limit, or timeout.

Multi-Role Session Tracking

Evidence is captured across separate, independently authenticated browser sessions for roles such as creator, approver, and supplier.

Route Memory

Successful runs distill their paths into step intents, providing advisory guidance for future executions without relying on fragile coordinates.

The Evidence Generation Workflow

1The system initializes a session using deterministic login scripts and encrypted secret stores.
2The AI agent executes versioned YAML objectives, recording screenshots and reasoning for every action.
3An independent verifier checks the results against defined assertions to determine the outcome.
4If a failure occurs, the system classifies the error and, if a product bug is suspected, flags a report for human review.
5All model calls are metered, and the final evidence is aggregated from the call level up to the release suite.

Frequently Asked Questions

Does the AI automatically file bug reports in the tracker?

No. To prevent false positives, the AI flags suspected product bugs with a ready-to-file report. A human must review and file it; the AI never asserts a defect on its own authority.

How does the system handle different AI models and providers?

Argus is model-agnostic. Model selection is managed in the database and read per run, allowing changes to take effect without redeploying code. It probes provider endpoints for vision and coordinate support rather than assuming compatibility.

Is this a replacement for traditional scripting tools like Playwright?

It complements them. Deterministic scripts are used for stable, predictable flows, while AI agents are deployed for long, multi-role, and frequently changing scenarios that are difficult to script.

How is the platform deployed and where is data stored?

The platform is self-hosted using ten containers and a shared artifact volume across its three engines (QA, AI, and pentest), ensuring there is no external SaaS dependency.

How are costs managed when using multiple AI models?

The system uses a two-tier execution model where a cheaper model runs first, escalating to a stronger model only upon agent-side failure. Every call is metered and aggregated by step, scenario, and suite.

Modernize Your QA Evidence

Experience autonomous regression testing with Argus QA. Contact Allmaz to learn how our AI-driven platform can secure your enterprise web applications.

Request a demo