Burrowbox Blog
← All posts

Devin signs in once and remembers. Build a QA agent that tests behind your login with the Claude Agent SDK

· The Burrowbox team · 8 min read

#What Devin does with browser logins

Devin is Cognition's AI software engineer (Cognition). Most web apps worth testing sit behind a login, and Devin's docs describe how it gets past one without asking every time (Devin docs):

  • A person logs in once through Devin's interactive browser, in the session's Desktop tab, and handles SSO redirects, MFA prompts and CAPTCHAs themselves. When they ask Devin to "save the browser profile", Devin zips the browser's cookies, localStorage and other Chrome profile data and attaches it to the organization's blueprint. Later sessions unpack it and start signed in.
  • Saved passwords and caches are left out of the profile. Restored cookies expire on the app's usual schedule, so the profile has to be saved again when the session runs out. Profiles are shared across the organization, and the docs recommend saving one for a dedicated service account rather than a personal one.

Credentials, site cookies and TOTP codes for two-factor logins can be stored separately as Secrets, at organization, personal or session scope (Devin docs).

Burrowbox isn't involved with Cognition or Devin.

#The pattern

A QA agent that logs in on every run spends time on the login page and adds a step that can break before the real test starts. And a model that types a password has that password in its context. Devin's approach handles both:

  1. The agent's browser starts signed in.
  2. It signs in as a test account made for the job.
  3. When it has to log in again, the secrets come from outside the prompt.

Devin stores the signed-in state once per organization, in the blueprint. On Burrowbox, the browser's state lives on the machine's disk, so it's stored once per machine. A machine is a Linux computer with one persistent browser, and stopping it keeps the browser's cookies (session cookies included) along with the rest of its files. The post on sandboxes and persistent machines covers what else survives a stop.

#Build it with the Claude Agent SDK and Burrowbox

The QA agent runs in your CI job with the Claude Agent SDK. It connects to a Burrowbox machine's MCP endpoint, works through a checklist in that machine's browser, and finishes with PASS or FAIL. The Claude computer use post wires up the same machine through the Messages API's MCP connector. This one uses the Agent SDK, which runs the tool loop for you and fits a script in CI.

#1. Create a machine for QA

Create the machine with "browser": "full" (Chromium). Both modes keep cookies across a stop and start, but the full browser also keeps sign-ins through live agent updates (Browser mode).

curl -X POST https://burrowbox.dev/api/machines \
  -H "Authorization: Bearer $BURROWBOX_KEY" -H "Content-Type: application/json" \
  -d '{"name": "qa-staging", "size": "small", "browser": "full", "ttlMinutes": 30}'
# → { "id": "…", "mcpUrl": "https://burrowbox.dev/api/machines/…/mcp", "mcpToken": "tmm_…" }

Save the id as QA_MACHINE_ID in your CI settings. With ttlMinutes, the machine stops by itself after 30 minutes and keeps its state.

#2. Put the staging test account in the vault

Each machine has an encrypted vault. The browser_login tool matches a credential to the page by URL, fills in the form, submits it and enters the TOTP code if the account has one. The password never comes back to the model, and vault_list returns only names, URLs and usernames.

curl -X PUT https://burrowbox.dev/api/machines/$QA_MACHINE_ID/vault/staging \
  -H "Authorization: Bearer $BURROWBOX_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://staging.example.com/login", "username": "qa-bot@example.com", "password": "…", "totpSecret": "JBSWY3DPEHPK3PXP"}'

After the first login, the cookies stay on the machine. Most runs start signed in and never touch the vault. When the session expires, the agent calls browser_login and carries on.

#3. The QA agent

npm i @anthropic-ai/claude-agent-sdk and npm i -D tsx. The SDK reads ANTHROPIC_API_KEY from the environment of the process that runs it (quickstart).

// qa.ts — run with: npx tsx qa.ts
import { query } from "@anthropic-ai/claude-agent-sdk";

const BB = "https://burrowbox.dev";
const id = process.env.QA_MACHINE_ID!;
const headers = { Authorization: `Bearer ${process.env.BURROWBOX_KEY}`, "Content-Type": "application/json" };
const bb = (path: string, init: RequestInit = {}) => fetch(`${BB}${path}`, { ...init, headers }).then((r) => r.json());

const CHECKLIST = `
1. Open https://staging.example.com/dashboard. If you land on the sign-in page, call browser_login, then continue.
2. The dashboard lists at least one project.
3. Settings > Billing opens without an error message.
4. Create a project named "qa-${Date.now()}" and check that it appears in the project list.
`;

// Start the machine (it may be stopped since the last run). The 30-minute deadline is a safety net.
await bb(`/api/machines/${id}/start`, { method: "POST", body: JSON.stringify({ ttlMinutes: 30 }) });
let machine = await bb(`/api/machines/${id}`); // includes mcpUrl and mcpToken
while (machine.status !== "running") {
  if (machine.status === "error") throw new Error(`Machine failed: ${machine.lastError}`);
  await new Promise((r) => setTimeout(r, 2000));
  machine = await bb(`/api/machines/${id}`);
}

const tools = ["browser_navigate", "browser_snapshot", "browser_click", "browser_fill", "browser_screenshot", "browser_login"];
let verdict = "FAIL";
try {
  for await (const message of query({
    prompt: `You are a QA tester. Work through this checklist in the browser. Use browser_snapshot to read each page. ` +
      `When a step fails, take a browser_screenshot and describe what you see. ` +
      `Report each step as PASS or FAIL with one line of evidence. End with a final line that is exactly PASS or FAIL.\n${CHECKLIST}`,
    options: {
      mcpServers: {
        burrowbox: { type: "http", url: machine.mcpUrl, headers: { Authorization: `Bearer ${machine.mcpToken}` } },
      },
      allowedTools: tools.map((t) => `mcp__burrowbox__${t}`),
      permissionMode: "dontAsk", // anything not listed above is denied
      maxTurns: 60,
    },
  })) {
    if (message.type === "result") {
      const report = message.subtype === "success" ? message.result : `Agent stopped: ${message.subtype}`;
      console.log(report);
      verdict = report.trim().split("\n").pop()!.trim();
    }
  }
} finally {
  await bb(`/api/machines/${id}/stop`, { method: "POST" }); // cookies and files are kept for the next run
}
process.exit(verdict === "PASS" ? 0 : 1);

Some notes on the script:

  • Only the tools it needs. allowedTools approves six browser tools by their mcp__<server>__<tool> names. With permissionMode: "dontAsk", any other call that would need approval is denied instead of waiting for a prompt that nobody in CI will answer (permissions).
  • A narrow token. The machine token (tmm_…) works only on this machine's MCP endpoint. It can't list, create, stop or bill anything (Authentication). Your API key stays in the script and never reaches the model.
  • Stop in finally. A failed run still turns the machine off, so it doesn't keep billing between deploys.

#Run it after every deploy

Add a job that runs after your staging deploy. The script's exit code fails the build when the agent reports FAIL.

  qa:
    needs: deploy-staging
    runs-on: ubuntu-latest
    concurrency: qa-staging   # one browser, so one run at a time
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22 }
      - run: npm ci
      - run: npx tsx qa.ts
        env:
          BURROWBOX_KEY: ${{ secrets.BURROWBOX_KEY }}
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          QA_MACHINE_ID: ${{ vars.QA_MACHINE_ID }}

The machine has one browser, so runs share its tabs and cookies. The concurrency group keeps two deploys from testing at the same time. If you need parallel runs, give each one its own machine and its own vault entry.

If you'd rather not run the agent in CI, Burrowbox can run a job inside the machine itself: a webhook with stopAfter: true wakes the machine, runs one action and stops it again, and an event webhook reports the result as run.succeeded or run.failed. The Comet post shows that wiring with a schedule.

#When the login needs a person

Some logins can't come from a vault: SSO with a push approval, or a CAPTCHA on the identity provider. As in Devin's setup, a person signs in once. Create an interactive live-view link with "mode": "browser" and "interactive": true and a short ttlSeconds, then open it and sign in as the test account. The cookies stay on the machine for later runs until the app expires them.

#What it costs

A small machine costs $0.11 an hour while running, billed by the minute, so a ten-minute QA run comes to about two cents of machine time. Stopped between deploys, it costs $0.001 an hour, about $0.72 a month. Model usage is billed separately by Anthropic. See Billing.

Create an account to try it.

#Key takeaways

  • Devin keeps a signed-in browser profile per organization. A Burrowbox machine keeps its browser's cookies on its own disk, once per machine.
  • Use a dedicated test account, keep its password and TOTP secret in the vault, and let browser_login handle expired sessions.
  • Give the agent only the machine token and an explicit tool list, and stop the machine when the run ends.

#FAQ

#How does Devin stay logged in to websites?

A person logs in through Devin's interactive browser, then asks Devin to save the browser profile. Devin adds the cookies, localStorage and Chrome profile data to the organization's blueprint, and new sessions start from it. Saved passwords aren't included, and cookies still expire normally.

#Does the QA agent see the test account's password?

No. browser_login fills and submits the form from the vault and returns a result, not the secret. vault_list shows names, URLs and usernames only.

The agent lands on the sign-in page, calls browser_login, and the vault credential (with its TOTP code, if set) signs it in again. If the login needs a person, for example an SSO push, open an interactive live-view link and sign in yourself.

#Can I use Python instead of TypeScript?

Yes. The Python Agent SDK takes the same server config through ClaudeAgentOptions(mcp_servers=…, allowed_tools=…), with "type": "http", the machine's mcpUrl and an Authorization header (Agent SDK MCP docs).

#Sources