Skip to main content
This example adds a pentest to your production deployment workflow. You’ll start with a script that runs tCell against your application, then connect it to GitHub Actions and limit testing to once every 24 hours. The check runs after deployment. Critical vulnerabilities fail the check so your team can investigate. The sections below explain each part. The complete files at the end combine them into a script and workflow you can copy into your repository.

1. Start a pentest

First, choose the production hostname you want to test and submit it for approval. This example checks for an exact hostname, such as api.example.com. Submit that hostname separately if you already have a wildcard target. Install the SDK and tsx, which runs the TypeScript script:
The script reads your API key and target from environment variables. Before launching an agent, it checks that both are present and the target is approved:
Retrieve tCell and give it a task containing the approved target:
At this point, testing has started in a sandbox. The Run object lets your script follow the work and retrieve its results.

2. Give the agent more context

You can customize tCell before starting the run. For this example, ask it to investigate access between accounts and give it constraints for testing production. Replace the agent lookup above with this configuration:
The spreads preserve tCell’s configuration and existing skills. The additional skill describes what to investigate, while the guardrail constrains how the agent should work. Use these fields for your application’s own procedures and constraints. See Skills and Guardrails for more examples. The call to agent.run() stays the same.

3. Follow the run and collect results

A pentest can take time. Iterate over the run to print activity as it happens, so someone viewing the CI logs can follow its progress:
run.wait() resolves when testing completes and throws if the run fails or stops. Once it resolves, retrieve the vulnerabilities:
You now have the vulnerabilities from the completed run. The next step is deciding how they affect the CI check.

4. Make the results a CI check

This example fails the check when it finds a critical vulnerability. It saves all vulnerabilities to a JSON file so your team can review the results, including issues below that threshold. Add the filesystem import at the top of the script, then save the results after retrieving them:
Set a nonzero exit code when a critical vulnerability is present:
The complete script wraps this work in main() and handles errors separately: It also saves run.json as soon as the run starts. That file contains the run ID, target, and triggering deployment’s commit, which you can use to find the run later.

5. Run after a production deployment

With the script in place, connect it to GitHub Actions. Add these values under your repository’s Settings → Secrets and variables → Actions: The workflow listens for deployment status updates and starts its job when the deployment succeeds in production:
Your deployment provider must report status through GitHub’s deployment status API. Change production if it uses a different environment name. After checking out your default branch and installing dependencies, the workflow passes the configured values to the script:
The final step uploads pentest-results/ as an artifact, including when critical vulnerabilities fail the check. Your team can download it from the workflow execution or review the vulnerabilities and evidence in Antigen. If deployment itself runs in GitHub Actions, you can instead add the pentest job after your deployment job with needs: deploy. Adapt the condition to that workflow and pass the deployed commit as DEPLOYMENT_SHA. Events created using a workflow’s GITHUB_TOKEN generally do not trigger another workflow. See Triggering a workflow.

6. Limit testing to once every 24 hours

Your team may deploy several times a day. Before installing dependencies or starting a pentest, the workflow checks whether its Run pentest step has already started in the last 24 hours. GitHub keeps that history, so the check can read earlier executions and their job steps. It includes previous attempts of rerun workflows. Once it finds a recent pentest attempt, it returns early:
The check runs in a step named cooldown. Later steps use its result to decide whether to proceed:
A skipped step does not extend the window. For example, if testing starts at 10:00 on Monday, deployments later that day skip testing. The first deployment more than 24 hours after that attempt can start another pentest. Failed attempts count too, so repeated deployments do not repeatedly launch testing after an error. A concurrency group makes subsequent jobs wait for the current job to finish. The next job checks the history before deciding whether to test:
This example shares one cooldown across the workflow’s production target. Keep the workflow, job, and pentest step names stable, and retain their history. If the history request fails, the job fails before testing starts.

Complete files

Save these files in your repository and commit them with package.json and package-lock.json to your default branch.
Testing runs against the live application, which may change during the pentest. The deployment SHA in run.json identifies what triggered testing. If CI is cancelled or times out, the agent can continue working. Use the run ID in the logs or artifact to retrieve or stop it.