Prerequisites
Prerequisites
- A test suite in Vitest, Playwright, or pytest
- A system under test whose webhook URL and recipient address you can set per test
- A DevHelm API key and workspace ID:
DEVHELM_WORKSPACE_ID and takes the key in its constructor. With curl, send the workspace as a header:Plan the isolation
Parallel jobs that share an inbox see each other’s events, and an old event can satisfy a new test’s wait. Isolate at the level each resource allows:
Inboxes in use at once are matrix shards × concurrent workflow runs, plus any that developers keep. Over the cap, the create step fails with
403 INBOUND_WEBHOOK_INBOX_LIMIT_REACHED.
Store the secrets
Create an API key used only by CI, so you can revoke it on its own (see API keys). Add it and the workspace ID as repository secrets:DEVHELM_WORKSPACE_ID even with a single workspace, and pass both on every step that calls DevHelm. Without it the CLI and SDKs call workspace 1 and fail with 404 “Workspace doesn’t exist”.
The workflow
Create.github/workflows/e2e.yml:
- The
ci-name prefix with run ID, attempt, and shard lets a sweeper find leftovers. - The URL is masked before it reaches
$GITHUB_ENV: anyone with it can write to the inbox, and logs on public repositories are public. - The key is set per step, so
npm ciand its install scripts never see it. if: always()runs the delete when tests fail or the run is cancelled. Deleting an inbox also deletes its events and raw bodies.
npx playwright test --shard=${{ matrix.shard }}/2; for pytest, pytest or pytest -n auto. The create and delete steps stay the same.
Write the tests
Each test builds its own path, registers<inbox URL><path> with the system under test, takes since before the trigger, and waits with pathPrefix. app-client stands for your own code.
timeoutMs. A longer runner timeout does not extend the HTTP call, and a shorter one kills the test before DevHelm answers. In pytest the timeout marker needs pytest-timeout.
Email: one address per test
address() returns a fresh address on the workspace’s assigned domain. Clear it when the test ends. There is no DevHelm Playwright package: a Playwright test calls the SDK the same way a Vitest test does.
signup.spec.ts
client.email.address(label=...) and calls clear() after the test. When CI has no mail provider, a test can inject the message with address.receive() instead; both paths run the same parsing. The wait response has otp and links but null text and html. The full flow, including links, is in Test a sign-up email flow.
Teardown
Never delete the assigned domain in teardown. Every concurrent run uses it, the next
address() call would create a domain with a new name, and on Free or Developer a second domain fails with 403 INBOUND_EMAIL_DOMAIN_LIMIT_REACHED.
Retention does not clean up for you. DevHelm does not delete captured data on a schedule today, so events and messages stay until teardown removes them. See Security and data.
Sweep inboxes from killed runs
A lost or timed-out runner skips the delete step, and its inbox keeps a slot. Run a sweeper on a schedule or at the start of each job, with a cutoff longer than your longest job:sweep-inboxes.ts
Parallel runs and the wait budget
All wait calls in an organization share 40 per minute: webhook and email waits, every workspace, every CI job, and every developer running tests locally. Over the limit, a wait returns429 RATE_LIMITED with Retry-After. A suite with 30 waits fits in one minute; two shards of it plus a second PR in flight use 120 and get throttled.
- Wait once per test, then assert everything on the returned event or message.
- Use one long
timeoutMs(up to 120000), not a loop of short waits. Each call counts. - Limit concurrent runs with a GitHub
concurrencygroup when the inbox cap or the wait budget is tight:
Back off on 429
Neither SDK retries on its own. Wrap waits in a helper that sleeps forretryAfter and tries again. Retrying with the same since is safe, because wait does not consume what it returns.
DevhelmRateLimitError and sleep for err.retry_after the same way. Size the runner timeout for the retries too: attempts × (timeoutMs + retryAfter).
Monthly quota
Every captured webhook request and every received or injected email counts toward one monthly event cap per organization: 300 on Free, 3,000 on Developer, 100,000 on Team. A suite that sends 30 events and runs 20 times a day needs 18,000 a month. At the cap, webhook senders and inject get403 INBOUND_EVENTS_MONTHLY_LIMIT_REACHED and nothing is stored. The counter updates asynchronously, so a burst can go slightly over first, and an inject accepted right at the cap can be dropped, which makes its wait return 408. See Limits, quotas, and errors.
Troubleshooting
Create step fails with 403 INBOUND_WEBHOOK_INBOX_LIMIT_REACHED
Create step fails with 403 INBOUND_WEBHOOK_INBOX_LIMIT_REACHED
The organization is at its inbox cap, counting every workspace. Delete leftovers from killed runs (the sweeper does this), reduce matrix shards, or add a
concurrency group.Wait returns 408 in CI but passes locally
Wait returns 408 in CI but passes locally
sincewas taken after the trigger, or not passed, and a fast event arrived before the wait started.- The test’s
pathPrefixdoes not match the path the system under test called. Paths are case-sensitive. - The system under test is slower in CI. Raise
timeoutMsand the runner timeout with it. - The monthly cap was reached. Check the sender’s response for
403.
The job is skipped on pull requests
The job is skipped on pull requests
The PR comes from a fork, and fork runs get no secrets. Run the suite after merge, or have a maintainer push the branch to the main repository.
Next steps
Wait for a webhook
Wait semantics, the Nth event, and negative assertions.
Wait for an email
Addresses,
subjectContains, and repeat sends.Limits, quotas, and errors
Every limit and error code with its fix.
Security and data
Keys in CI, captured headers, and cleanup.