# Connect Flock to your assistant.

Recommend focused interface tests, launch approved runs, and retrieve evidence-backed results from the assistant you already use.

**Flock MCP endpoint:** `https://mcp.flocksynthetics.com/mcp`

## Setup

### Codex

Put this native Codex MCP configuration in ~/.codex/config.toml; it uses the preregistered local callback.

```toml
[mcp_servers.flock-prod]
url = "https://mcp.flocksynthetics.com/mcp"
startup_timeout_sec = 350
tool_timeout_sec = 350
required = false

[mcp_servers.flock-prod.oauth]
client_id = "flock-codex"
callback_url = "http://localhost:8765/callback"
callback_port = 8765
```

Read-only login:

```sh
codex mcp login flock-prod --scopes "https://mcp.flocksynthetics.com/mcp/read"
```

Read + test opt-in:

```sh
codex mcp login flock-prod --scopes "https://mcp.flocksynthetics.com/mcp/read,https://mcp.flocksynthetics.com/mcp/test"
```

Leave oauth_resource unset. Use the first command for read-only access; use the second only when you explicitly want launch, retry, and cancel actions. Keep required = false for normal setup; use the per-run -c mcp_servers.flock-prod.required=true override only for a deliberate cold-start acceptance check. Keep this environment in its own flock-prod entry; do not overwrite another Flock environment's connection.

### Claude Code

Add one HTTP server entry, then let Claude Code open the fixed local callback.

Read-only setup:

```sh
claude mcp add-json flock-prod \
  '{"type":"http","url":"https://mcp.flocksynthetics.com/mcp","oauth":{"clientId":"flock-claude","callbackPort":8766,"scopes":"https://mcp.flocksynthetics.com/mcp/read"}}'
claude mcp login flock-prod --no-browser
```

Read + test opt-in:

```sh
claude mcp add-json flock-prod \
  '{"type":"http","url":"https://mcp.flocksynthetics.com/mcp","oauth":{"clientId":"flock-claude","callbackPort":8766,"scopes":"https://mcp.flocksynthetics.com/mcp/read https://mcp.flocksynthetics.com/mcp/test"}}'
claude mcp login flock-prod --no-browser
```

The scopes value is space-separated. If flock-prod already exists, update that entry rather than adding a duplicate. Keep each Flock environment in a distinct entry. Scope changes require disconnecting the old grant in Flock Settings → Connected tools and consenting again.

### Other clients

Use Streamable HTTP OAuth with PKCE and a preregistered client ID and callback.

```text
server URL: https://mcp.flocksynthetics.com/mcp
read scope: https://mcp.flocksynthetics.com/mcp/read
test scope (opt-in): https://mcp.flocksynthetics.com/mcp/test
```

Automatic arbitrary dynamic client registration is not supported. Flock does not issue a secret or API key for this connection.

## Workflow

1. Connect and authorize the requested read or test scope.
2. Confirm the account and choose the workspace to use.
3. Describe the application URL and the user journey; ask for a recommendation.
4. Review readiness and the estimate, then explicitly approve a launch.
5. Launch with a unique caller idempotency key and retain the receipt.
6. Follow lifecycle status, then retrieve results and inspect available evidence images.

## Agent guide

The 14 tools are `account`, `workspaces`, `list_auth_profiles`, `recommend_test_configuration`, `estimate_run`, `launch_run`, `get_run_status`, `get_run_results`, `get_run_image`, `retry_run`, `cancel_run`, `list_reviews`, `get_latest_review`, `get_review`.

- **account, workspaces:** Confirm the signed-in account and selected workspace.
- **list_auth_profiles:** Check a saved sign-in profile; never paste passwords, tokens, cookies, or browser storage into chat.
- **recommend_test_configuration, estimate_run:** Suggest a context and catalog based on your URL and goal, then explain readiness and expected usage.
- **launch_run, get_run_status, get_run_results:** Run an explicitly approved test, follow its batch lifecycle, and retrieve actionable findings plus bounded image descriptors.
- **get_run_image:** Read one actual run image from an opaque descriptor returned by get_run_results, with safe attribution metadata.
- **retry_run, cancel_run:** Retry or cancel an accepted run when the test scope and workspace policy allow it.
- **list_reviews, get_latest_review, get_review:** Read workspace reviews with bounded findings and run references; use get_run_results for a returned run.

Tool boundaries:

- `recommend_test_configuration`: Call with workspace_id, url, and goal. It is a context and catalog heuristic; it does not inspect the application.
- `estimate_run`: Call with workspace_id, url, and the proposed configuration. Do not add idempotency_key; its schema does not accept one.
- `launch_run`: Supply a unique caller idempotency_key. If the response is uncertain, replay the exact same payload and key.
- `retry_run`: Supply workspace_id, batch_id, and a unique caller idempotency_key for the retry. If the response is uncertain, replay the exact same payload and key.
- `get_run_status / get_run_results`: Status uses workspace_id and batch_id; results use workspace_id and run_id. Check lifecycle before treating results as final. Results report images_status as available, not_ready, or unavailable, with images_reason and images_retryable. Available images include opaque artifact_id, kind, label, content_type, size_bytes, origin_scope, and optional source_step.
- `get_run_image`: Use only workspace_id, run_id, and an artifact_id from get_run_results.images. Success returns native MCP image content plus structured status available and safe attribution metadata. Read failures are isError results with structured status not_ready or unavailable, reason, and retryable. Images are bounded to 8 MiB without truncation. Do not fetch arbitrary URLs or construct artifact IDs.
- `list_reviews`: Requires workspace_id. The default limit is 10 and the maximum is 50; source accepts all, dashboard, ci, or mcp. Page tokens are opaque, and each review_id is the canonical batch_id.
- `get_latest_review`: Requires workspace_id. Report clearly when the latest review is partial, degraded, or unavailable; do not imply freshness that the source cannot support.
- `get_review`: Requires workspace_id and review_id, which is the canonical batch_id. Details are bounded to at most 50 findings and 100 run references. Read lifecycle and technical findings before calling get_run_results for a returned run, and use repair_prompt_context.rendered_markdown as the canonical repair handoff when its status is available.

## Starter prompts

> Use Flock to recommend a small test of [URL] for [user journey]. Confirm my workspace, explain what you can assess, check saved sign-in readiness if needed, and show the estimate. Wait for my approval before launching.

> Use the read-only Flock connection to recommend and estimate a focused test. Do not launch anything.

> Investigate my failed review and give me what I need to fix it. Discover my account, workspace, latest review, and related runs; explain the lifecycle state before requesting final results; include technical findings and the canonical repair prompt; and inspect the available evidence images. Use read-only access and do not require me to supply IDs or know the dashboard structure.

## Review reads

> List reviews for workspace_id [WORKSPACE_ID] with the default limit of 10, source mcp, and opaque paging.

> Get the latest review for workspace_id [WORKSPACE_ID] and say plainly if the result is partial, degraded, or unavailable.

> Get review_id [BATCH_ID] for workspace_id [WORKSPACE_ID], summarize up to 50 findings and 100 run references, then use get_run_results for a returned run.

> Investigate my failed review without asking me for IDs. Discover the account and workspace, find the latest review, explain its lifecycle and technical findings, retrieve related run results only when ready, use the available canonical repair prompt, and inspect its evidence images.

## Permissions and safety

- **Read scope:** Account, workspaces, recommendations, estimates, status, results, workspace review reads, and manifest-bound run images.
- **Test scope:** Launch, retry, and cancel, subject to the selected workspace policy and spending controls.
- **Sign-in profiles:** Use a saved profile when a target requires login. Complete setup in Flock; never share credentials or browser secrets in chat.
- **Estimates:** An estimate describes expected usage. It is not a monetary guarantee or a claim that a run is free.

Use a saved auth profile for targets that require sign-in. Complete profile setup in the Flock dashboard; never put passwords, tokens, cookies, or browser storage in chat. Sign into Flock only when the assistant asks through authorization, and accept a workspace invitation first if the dashboard says **Invite required**.

If a launch response is interrupted, recover it with the original exact payload and idempotency key. Never choose a new key for recovery.

[HTML guide](/) · [Support](mailto:hello@flocksynthetics.com) · [Privacy](https://flocksynthetics.com/privacy/)
