This repository contains GUI test cases, bug injection scripts, and reproducible environments for four web applications. It is derived from the original WebTestPilot benchmark, with revised test steps and expected results. The change_* UI experiments are no longer part of this benchmark.
<app>/
test_cases/<task>.yaml # ordered test steps and expected results
bugs/<task>.js # optional bug injection script for the same task
environment/ # Compose, seed data, app images, and app configuration
runtime/ # shared browser, login, and bug injection files
app_config.json # app URLs, copied assets, and account reference data
template.yaml # YAML example
template.js # bug script example
There are 100 test cases across BookStack, Indico, Invoice Ninja, and PrestaShop. The YAML filename identifies the task; a bug script with the same filename stem is used when that task is run with bug injection enabled. Each app's environment/ contains its Compose stack and seed assets. Shared browser startup, login setup, and bug injection live in runtime/. A runner assembles these inputs into isolated task environments.
Each YAML file describes one task. steps is an ordered list: action tells the tester what to do, expectation states the expected visible result, and ground_truth optionally supplies a Playwright Python assertion for that result.
name: Comment
setup_function: login_to_bookstack
steps:
- action: From the dashboard click 'Page Template' link
expectation: Page contains title 'Page Template'
ground_truth: |
expect(page.locator("#bkmrk-page-title")).to_match_aria_snapshot("- heading \"Page Template\" [level=1]")
- action: Click 'Add Comment'
expectation: A WYSIWYG comment editor is open
ground_truth: |
expect(page.get_by_role("button", name="Save Comment")).to_be_visible()| Field | Meaning |
|---|---|
name |
Human-readable task name. |
setup_function |
Optional runner setup or login function to call before the task. |
description |
Optional description of the task. |
steps[].action |
Action to perform in the browser. |
steps[].expectation |
Expected result in natural language. |
steps[].ground_truth |
Optional Playwright Python assertion snippet. |
The filename stem is the stable task identifier. Use a unique stem within each app, and keep a matching bug script under bugs/ when testing bug detection. A bug script contains isConditionMet and onConditionMet blocks delimited by the markers shown in template.js.
These are the application image versions and web login accounts defined by this repository's environment files and login setup; database credentials are separate. The PrestaShop 8.2.8 Apache image is pinned by digest in its app Dockerfile.
| Web app | Application version / image | Login | Password | Notes |
|---|---|---|---|---|
| BookStack | solidnerd/bookstack:25.2.1 |
admin@admin.com |
password |
Admin account |
| Indico | 3.3.6 (pip install indico==3.3.6) |
admin@admin.com |
webtestpilot |
Admin account |
| Invoice Ninja | invoiceninja/invoiceninja-debian:5.11.61-d |
admin@admin.com |
password |
Admin account |
| PrestaShop | prestashop/prestashop:8.2.8-apache |
admin@admin.com |
admin12345 |
Seller; admin path /webtestpilot/ |
| PrestaShop | prestashop/prestashop:8.2.8-apache |
auto.customer@example.com |
mypassword |
Buyer account |
- Add
<app>/test_cases/<task>.yamlusingtemplate.yamland an existing test case as examples. Give each step an action and a specific expectation; addground_truthassertions where they can check the result directly. - Set
setup_functionwhen the task needs a particular logged-in session. The supported functions are registered inruntime/init.py. - If the task has an injected bug, add
<app>/bugs/<task>.jsusingtemplate.js. Keep the// BEGINand// ENDmarkers around both functions. - Run the task through the intended runner to check the setup, each step, and the bug condition.
- Create
<app>/test_cases/,<app>/environment/, and, if the app has injected bugs,<app>/bugs/. Add YAML cases and matching bug scripts using the existing naming convention. - Add the app's
docker-compose.yaml,seed.sql, andseed-loader.shunder<app>/environment/. Put app-specific Dockerfiles, config files, and optionalbaseline.sqlthere too. Seeruntime/BASELINE.mdfor baseline semantics. - Add an entry to
app_config.jsonwithapp_url,extra_files, andextra_dirsfor app-specific assets copied into each task. Itscredentialsfield is reference documentation; the login code reads its own values fromruntime/init.py. - Implement any new
setup_functioninruntime/init.py, register it in_SETUP_FUNCTIONS, and update the version and test account table above. - Generate and run the app's tasks with a consumer of this benchmark to verify its environment and steps.