GOOSY|Docs
← Home

Agentic testing

Runs, self-heal & monitoring

Run a suite, read the evidence, and let drifted tests fix themselves.

Running a suite

“Run now” queues a run and returns immediately — execution happens asynchronously, and the run's status updates from queued to running to a terminal state as it progresses. A suite runs its active tests in order against the chosen environment, respecting that environment's requests-per-second cap across the whole run, not per test.

A test whose spec changed since it was compiled (marked stale) is recompiled against the current spec automatically before it runs, so a run always reflects the current API shape rather than failing on a stale artifact for no useful reason.

Reading a result

OutcomeMeaning
passEvery step's assertions held.
failAn assertion failed and was classified as a real regression (or the repair agent couldn't tell).
repairedThe test drifted from a real API change; Goosy recompiled it and the recompiled version passed.
env_issueA timeout or network error, retried automatically before being recorded — never silently dropped.

Every step's evidence — request method, URL, headers, and body, and the response status, headers, and body — is stored with the result. Auth headers and any value matching a stored secret are redacted before evidence is ever persisted, in both directions (request and response), not just in headers named like an auth header.

Self-heal: how a failure gets classified

A test failure never just sits there unexplained. Goosy diffs the spec version the test was compiled against against the current spec, and returns exactly one verdict:

  • Drift — the API intentionally changed. Goosy recompiles the test against the current spec and reruns it once; if that passes, the result is recorded repaired with a note on what changed.
  • Regression — the spec is unchanged but the API broke. Recorded as a failure and surfaced prominently — this is the case that should get your attention first.
  • Flake — looks environment-related rather than a real contract issue. Retried a bounded number of times before being recorded.
  • Inconclusive — not enough signal to call it either way. The test is marked for review rather than silently disabled or silently trusted.

Needs attention

One queue aggregates everything that needs a human: tests that failed to compile, tests the repair agent couldn't classify, and recent regressions. Each entry can be dismissed (which never deletes the underlying test or its history) or edited and recompiled directly from the queue.

Dashboards

Per suite: average and p95 run latency, pass rate, repair rate (of tests that failed and got a repair attempt, how many the agent actually fixed), and average tokens spent per run — a passing suite with no drift costs zero tokens, since the run stage never calls a model.

Cost and safety by default
Every model call this module makes goes through the same token-budget governor every scan uses, with an additional hard ceiling per run — if a run's repair activity would exceed it, the run fails explicitly rather than silently truncating results. Runs also carry a hard per-run timeout, and a per-step request timeout independent of it.
← Environments, specs & suitesCLI →