Bring your live agents into one control panel. Track every run, understand the cost, and pause, restart, swap the model, push a prompt, or teach it a new skill—without redeploying the target app.
Compatibility depends on access to your app’s code or HTTP hooks. Provider usage and accounts stay with you.
03:07 · The operating gap
From the first signal to the next safe run.
When an agent loops at 3am, observability tells you what happened. ACP carries the journey through the missing half: isolate, intervene, and verify—while the app stays deployed.
Runs, token use, cost, errors, and health put the problem on one screen.
Observability tools already do this part well.
02Diagnose
Find the agent.
Separate the misbehaving agent from the rest of the app before changing anything.
No codebase hunt. No guessing which deploy is live.
03Intervene
Pause it now.
Send the command to the integration already running inside the target app.
The local wrapper checks state before the next LLM call.
04Verify
Watch the line flatten.
The next heartbeat confirms the new state. Production stays up; the model calls stop.
No prompt edit. No build queue. No redeploy.
The 3-minute connect
Your app. Your agents. One connection.
Create an app, install agent-control-panel or paste the generated instructions into your coding agent, and wait for the first heartbeat. Scroll to follow the connection, from your first key to a live heartbeat.
Your app reports upward on its own schedule. ACP sends signed commands downward. Nothing sits in front of your LLM traffic, so ACP being down cannot take your agents with it.
Upward · your app reports
Heartbeat every 30s
POST/api/heartbeat
POST/api/report
GET/api/prompts/:agentId
Status, tokens, runs and errors since the last beat. Run reports carry model used, latency and prompt version.
The control plane
ACP holds the state
commandspause · resume
commandsstop · restart
commandsupdate_model
commandsupdate_prompt
Six commands, one audit trail. Prompt pushes are approval-gated before they reach the app.
Downward · signed webhook
HMAC, timestamp, nonce
headerx-acp-signature
headerx-acp-timestamp
headerx-acp-nonce
Signature over ts.nonce.body, a 300-second freshness window, and a replay cache that rejects a reused nonce.
See + operate
See what happened. Decide what happens next.
Five surfaces, one control plane: send commands, read failing runs, account for tokens, score responses, and turn a repeated failure into a deployed skill.
The acknowledgement will appear here. This demo does not contact a real agent.
app.acp / runs1,284 runs today · 11 failed
run_8f21c4
checkout-recovery · gpt-4.1-mini · 6 tool calls
failed
Error
ToolTimeout: stripe.refund did not return within 30000ms
Why it failed
The refund tool blocked on a Stripe call that never returned. The agent re-planned after each timeout and repeated the same call six times inside one run, so it spent 18.4k tokens before giving up.
How to fix it
01Raise the tool timeout to 60s and send an idempotency key so a retry cannot double-refund.
02Cap in-run retries at two, then fail fast instead of re-planning.
03Push prompt v13, which tells the agent to stop after a tool error and hand off.
app.acp / token usage
Total
6.4M
Input
4.3M
Output
2.1M
Input vs output per dayinputoutput
Estimated cost$284.10
checkout-recoveryorchestrator · 41%2.6M
support-routerrouter · 24%1.5M
invoice-parsertool · 19%1.2M
digest-writeragent · 11%0.7M
lead-enricherworker · 5%0.4M
checkout-suite14 agents · 52%3.3M
support-desk9 agents · 26%1.7M
billing-ops6 agents · 15%0.9M
growth-loops2 agents · 7%0.5M
gpt-4.1-mini63% of tokens4.0M
claude-sonnet-4-629% of tokens1.9M
gpt-4.18% of tokens0.5M
app.acp / feedbackrun_8f21c4 · awaiting review
Agent response
“I was unable to process the refund. Please contact support during business hours and reference order 41822.”
Your score
not scored
Score the run, or let the reviewer model score it. Low scores collect against the agent and become the evidence for a new skill.
Reviewer model scoreauto
Resolution34
Policy48
Tone88
Overall57
3 of the last 20 runs scored under 60 on resolution.Draft a skill
01Failure detected11 failed runs, same tool timeout
02Skill draftedretry policy + handoff wording
03Tested in the lab40 replayed runs, agent A/B
04Deployed on passonly if the overall score clears the gate
Before / after · 40 replayed runs
Pass rate61%→94%
Repeat failures11→1
Tokens per run18.4k→4.2k
Review the results before applying measured: pendingnot deployed
01Surface 01 of 05
Control
Pause, restart, swap the model or push a prompt. The command lands in the integration inside your app.
ACP never routes your LLM traffic. Local state is checked before the next call, so nothing waits on a deploy. Pick a command to send it in the panel.
02Surface 02 of 05
Runs
Open the failing run, read the reason it failed, and get the fix in the same view.
Every run report carries model, latency, tokens, tool calls and the error. Failures are grouped so a repeated cause is obvious instead of buried in a log.
Filter to failures, one agent, or one app.
Read the error, the explanation of what the agent did, and the steps to fix it.
Act from the run: retry it, push a prompt, or send it to the Skills Hub.
When an agent keeps failing, it gets taught a skill instead of a patch.
ACP drafts the skill from the failing runs, replays those runs against it in the Skills Testing Lab, and compares the agent before and after. The skill deploys only if the overall result clears the gate.
Draft a skill from the failures that prove it is needed.
Test it in the lab on replayed runs, before and after, side by side.
Deploy on pass only. A skill that does not clear the gate stays a draft.
ACP brings run diagnostics, quality scoring, and replay evaluation together with live agent controls. Compare the workflows you need, from investigating a failure to testing a fix and applying an approved change.
Remote control is only acceptable if the wire is signed and the failure mode is boring. Both are enforced in the generated integration, not in documentation.
Compared with a timing-safe check, and verified against the current and previous secret so rotation never drops a command.
Replay defence
300-second window, nonce cache
A stale timestamp is rejected with 403. A nonce that has already been seen is rejected with 403.
Fails open
Reporting errors are caught
Heartbeat and report failures log and return. A prompt fetch failure falls back to cache, then to the hardcoded prompt.
Not in the path
No proxy, no key custody
Your provider keys stay in your app. ACP never sits between your agent and the model.
Objections
The questions a skeptical developer asks.
Answered against what the code actually does today.
No. The integration is a single generated file you paste into the app you already have. Agents stay hardcoded where they are; the wrapper reads their state before each call.
Your agents keep running. Heartbeats and run reports are wrapped in try/catch and only log on failure. Prompt fetches fall back to the local cache, then to the prompt already in your code.
No. There is no gateway and no proxy. A model swap is a local override the wrapper resolves before your own request goes out, so your provider keys never leave your app.
The webhook sets paused state immediately, and the next wrapped run returns before calling the model. The dashboard confirms it on the following heartbeat, within 30 seconds.
No. Deep tracing and evaluation are stronger elsewhere, and the comparison table above says so. Keep them for analysis and use ACP for the intervention they cannot perform.
Zero. The system ran for a year as one operator's production control plane for 320 agents across 31 applications, and the hosted product is at the start of external validation. You would be early.
0agents operated
0applications
1 yearin daily production
Who built it
Built from the operator’s seat.
The founder designed, built, and operated the system alone for a year as the real control plane for a crypto venture fund. It managed 320 agents across 31 applications in daily production.
That history proves the operating model. ACP currently has zero external connected apps; the public product is at the beginning of that validation.
The whole evaluation
Connect one app. Operate one real agent.
If you can see a heartbeat and pause a live agent without a deploy, you understand the product.
1Create the appSign up and receive an API key and webhook secret. No card, no sales call.
2Paste the integrationHand the generated file to your coding agent, or drop it in yourself.
3Take controlWatch the first heartbeat land, then pause the agent from the dashboard.