Glossary · Argus QA

What is a flaky test?

What is a flaky test? A clear explanation for Azerbaijani business — and how Argus QA applies it.

Solving the Challenge of Flaky Tests

A flaky test is a software test that yields both passing and failing results across multiple executions without any changes to the underlying code. This inconsistency often stems from unstable environments, timing issues, or fragile selectors, making it difficult for teams to distinguish between a genuine product defect and a testing artifact. In enterprise web applications, this instability often leads to 'test fatigue,' where developers ignore failures due to a lack of trust in the testing suite. Argus QA addresses this by shifting from brittle, coordinate-based scripts to autonomous regression testing using AI computer-use agents. By utilizing versioned YAML objectives rather than static selectors, the system focuses on intent and outcomes. This approach allows the testing engine to observe screenshots, reason through actions, and navigate complex, multi-role workflows in a way that mimics human interaction, significantly reducing the noise associated with traditional automated testing.

Capabilities

Enterprise-Grade Stability and Reliability

Elimination of fragile selectors and coordinates through objective-based YAML definitions

Precise failure classification to distinguish between product defects, agent errors, and environment issues

Prevention of data collisions in shared environments via run-scoped test data naming

Removal of agent bias through a dedicated independent verifier for all assertions

Secure, deterministic authentication using scripts and an encrypted secret store

Optimized cost and reliability via a two-tier model execution and escalation strategy

Core Capabilities of Argus QA

Objective-Based Scenarios

Scenarios are defined as versioned YAML objectives containing roles, assertions, and step/time budgets, removing the need for brittle selectors.

Granular Failure Classification

Failures are categorized as product defects, agent failures, environment failures, assertion failures, loops, step limits, or timeouts.

Independent Assertion Verification

To ensure objectivity, assertions are checked by an independent verifier rather than the agent that performed the action.

Route Memory Guidance

Successful runs distill their paths into step intents, providing advisory guidance for future runs without relying on fixed coordinates.

Deterministic Login Flow

Authentication is handled by deterministic scripts and an encrypted secret store, ensuring the agent never fails during the login phase.

The Autonomous Testing Workflow

1The agent observes a screenshot and states its reasoning based on the YAML objectives.
2One action is taken per turn, with every step persisted regardless of whether it passed or failed.
3Multi-role workflows execute in fresh, independently authenticated browser sessions for each role.
4A cost-effective model handles initial execution, escalating to a stronger model only for agent-side failures.
5An independent verifier checks assertions to determine if the objective was successfully met.
6Suspected product bugs are flagged with a ready-to-file report for human review and final filing.

Frequently Asked Questions

Does Argus QA replace traditional testing frameworks like Playwright?

No, it complements them. Use deterministic scripts for stable, unchanging flows and AI agents for long, multi-role, or frequently changing scenarios.

How does the system prevent data collisions during parallel execution?

Argus QA utilizes run-scoped test data naming, which allows multiple parallel agents to operate against a shared environment without interfering with one another.

Can the AI autonomously file bugs in our tracking system?

No. To maintain accuracy, the AI flags suspected bugs and generates a report, but a human must review and file the defect on their own authority.

How is the cost of AI model usage tracked and managed?

Every model call is metered. Costs are aggregated from the individual call level up through the step, scenario, suite, and finally the release.

Is Argus QA a SaaS product or self-hosted?

It is entirely self-hosted, consisting of ten containers and a shared artifact volume, ensuring no external SaaS dependency.

Stabilize Your QA Pipeline

Move beyond flaky tests with Argus QA's self-hosted AI testing platform. Contact Allmaz to learn more about our autonomous regression testing.

Request a demo