THE VERIFICATION LAYER FOR YOUR SOFTWARE FACTORY
Black-box tests for your web apps, written and maintained as code by AI agents, and run on infra that lets you ship fast.
Product walkthrough
Trusted by fast growing teams
01 · Black box
Tests that only see what users see
Our tests don't know or care about your source code. They drive a real browser against your app (click, type, and check what's on the screen) so they stay put when the code underneath changes. Migrating to Rust? Your tests don't know, and don't change.
- No hooks, mocks, or test-only backdoors in your app
- Refactors and rewrites don't break the suite
- Any stack, any framework, any language
148 files changed
await page.getByRole("button", { name: "Pay now" }).click();
await expect(page.getByText("Order confirmed")).toBeVisible(); ✓ 212 passed · before and after
02 · Infra
Fast infra for browser tests
Playwright on CI splits tests by file order, so one slow shard holds up the whole run. We schedule every test from its duration history, longest first, so shards finish together.
- Duration-aware sharding, retries, and parallel workers
- No CI config or runners to maintain
- Video, trace, and logs for every test
WALL CLOCK: 21 MIN
WALL CLOCK: 17 MIN
03 · Agent
A suite that maintains itself
When your app changes, the agent figures out whether a failure is a bug, a flake, or a test that drifted, then fixes the test or tells you about the bug, in Slack, with evidence.
- Writes new tests from a prompt, a PR, or a ticket
- Triages every failure and repairs drifted tests
- Every change lands as a PR, reviewed by an engineer
- 09:12 DEPLOY checkout v2.4 shipped to staging
- 09:31 RUN 3 failures in checkout/
- 09:33 AGENT Triaged: 2 tests drifted (button renamed), 1 real regression
- 09:34 AGENT Posted to #releases: coupon field rejects valid codes
- 09:41 AGENT Opened PR #1260: updated 2 tests, re-ran green
- 09:58 HUMAN Reviewed and merged
04 · Composable
Primitives you can build on
Everything in the dashboard is a primitive with an API and a CLI. Trigger runs from your pipeline, hand work to the agent from your own coding agent, and pull results wherever you need them.
- CLI and API for every action in the product
- Agent skill for Claude Code, Codex, and friends
- Results in GitHub checks, Slack, and your tracker
$ empirical session -x "add a test for coupon codes at checkout" ✓ PR #1262 opened · 1 test added · run passed $ empirical api api/test-runs/4821/status { "status": "failed", "passed": 211, "failed": 1 }
Stories and writing
- Your agents don't belong in your codebase
- Writing E2E tests is the easy part. Keeping them green is what kills you.
- Unit tests didn't catch what our coding agents are breaking.
- 500+ Playwright tests on every PR, no QA team
- videostil: Bringing Video Understanding to Every LLM
- Appwright: A new test framework for mobile apps






