Control plane for live agents

Operate the agents you already shipped.

Bring your live agents into one control panel. Track every run, understand the cost, and pause, restart, swap the model, push a prompt, or teach it a new skill—without redeploying the target app.

No framework adoption · No LLM proxy · Integration fails open

acp / control plane live
Topology Anomaly isolated to agt_84c1
Spend today $42.18
+184% anomalycommand ready
agentcheckout-recovery
statusrunning
modelgpt-4.1-mini
promptv12
Simulated
Your stack. One operating view.

Built around the tools you already use.

Keep your models and your builders. Connect ACP through the SDK or HTTP integration inside your app.

Model providers

OpenAIAnthropicGeminiGroqDeepSeekMistral AIMeta LlamaOllamaOpenRouterCerebrasQwenKimiMiniMax AIZ.ai / GLMxAIPerplexityHugging FaceNVIDIAIBMMicrosoftMetaThinking Machines

App builders

ReplitLovableBolt.newBase44BubbleFirebase Studiov0

Agent & workflow tools

LangChainLangGraphCrewAILangflown8nMakeZapierLindyManusSalesforce Agentforce

Compatibility depends on access to your app’s code or HTTP hooks. Provider usage and accounts stay with you.

03:07 · The operating gap

From the first signal
to the next safe run.

When an agent loops at 3am, observability tells you what happened. ACP carries the journey through the missing half: isolate, intervene, and verify—while the app stays deployed.

Incident / cost anomaly 03:07:04

Spend accelerating

checkout-recovery · agt_84c1

anomaly
command sent
Retries 21
Cause unresolved
Next step inspect
03:07:49POST /api/agents/:id/commands · pausesent
03:08:04webhook acknowledged · x-acp-signature okok
03:08:34next heartbeat · status pausedverified
01 Detect

See the spike.

Runs, token use, cost, errors, and health put the problem on one screen.

Observability tools already do this part well.
02 Diagnose

Find the agent.

Separate the misbehaving agent from the rest of the app before changing anything.

No codebase hunt. No guessing which deploy is live.
03 Intervene

Pause it now.

Send the command to the integration already running inside the target app.

The local wrapper checks state before the next LLM call.
04 Verify

Watch the line flatten.

The next heartbeat confirms the new state. Production stays up; the model calls stop.

No prompt edit. No build queue. No redeploy.
The 3-minute connect

Your app. Your agents.
One connection.

Create an app, install agent-control-panel or paste the generated instructions into your coding agent, and wait for the first heartbeat. Scroll to follow the connection, from your first key to a live heartbeat.

Connection lab / production API credentials ready
Generated environment variables
ACP_URLhttps://acp.skaigroup.techpublic
ACP_API_KEYacp_••••••••••••••••••••••••one time
ACP_AGENT_IDagt_84c1per agent
ACP_WEBHOOK_SECRET••••••••••••••••••••••••••••secret
Keys are generated by ACP. Store them in the target app's environment.
How it works

Two directions, one wire.

Your app reports upward on its own schedule. ACP sends signed commands downward. Nothing sits in front of your LLM traffic, so ACP being down cannot take your agents with it.

Upward · your app reports

Heartbeat every 30s

POST/api/heartbeat
POST/api/report
GET/api/prompts/:agentId

Status, tokens, runs and errors since the last beat. Run reports carry model used, latency and prompt version.

The control plane

ACP holds the state

commandspause · resume
commandsstop · restart
commandsupdate_model
commandsupdate_prompt

Six commands, one audit trail. Prompt pushes are approval-gated before they reach the app.

Downward · signed webhook

HMAC, timestamp, nonce

headerx-acp-signature
headerx-acp-timestamp
headerx-acp-nonce

Signature over ts.nonce.body, a 300-second freshness window, and a replay cache that rejects a reused nonce.

See + operate

See what happened.
Decide what happens next.

Five surfaces, one control plane: send commands, read failing runs, account for tokens, score responses, and turn a repeated failure into a deployed skill.

Control Runs Token usage Feedback Skills Hub
Remote command Interactive simulation

checkout-recovery

agt_84c1 · production

running
Pause agentsigned webhook v2
commandpause
targetagt_84c1
deploynot required
Changes apply in the target app.
Ready to send

The acknowledgement will appear here. This demo does not contact a real agent.

app.acp / runs1,284 runs today · 11 failed

run_8f21c4

checkout-recovery · gpt-4.1-mini · 6 tool calls

failed
Error

ToolTimeout: stripe.refund did not return within 30000ms

Why it failed

The refund tool blocked on a Stripe call that never returned. The agent re-planned after each timeout and repeated the same call six times inside one run, so it spent 18.4k tokens before giving up.

How to fix it
01Raise the tool timeout to 60s and send an idempotency key so a retry cannot double-refund.
02Cap in-run retries at two, then fail fast instead of re-planning.
03Push prompt v13, which tells the agent to stop after a tool error and hand off.
app.acp / token usage
Total

6.4M

Input

4.3M

Output

2.1M

Input vs output per day inputoutput
Estimated cost $284.10
checkout-recoveryorchestrator · 41% 2.6M
support-routerrouter · 24% 1.5M
invoice-parsertool · 19% 1.2M
digest-writeragent · 11% 0.7M
lead-enricherworker · 5% 0.4M
checkout-suite14 agents · 52% 3.3M
support-desk9 agents · 26% 1.7M
billing-ops6 agents · 15% 0.9M
growth-loops2 agents · 7% 0.5M
gpt-4.1-mini63% of tokens 4.0M
claude-sonnet-4-629% of tokens 1.9M
gpt-4.18% of tokens 0.5M
app.acp / feedbackrun_8f21c4 · awaiting review
Agent response

“I was unable to process the refund. Please contact support during business hours and reference order 41822.”

Your score

not scored

Score the run, or let the reviewer model score it. Low scores collect against the agent and become the evidence for a new skill.
Reviewer model score auto
Resolution34
Policy48
Tone88
Overall57
3 of the last 20 runs scored under 60 on resolution. Draft a skill
app.acp / skills hub · testing labskill_refund_timeout draft
01Failure detected11 failed runs, same tool timeout
02Skill draftedretry policy + handoff wording
03Tested in the lab40 replayed runs, agent A/B
04Deployed on passonly if the overall score clears the gate
Before / after · 40 replayed runs
Pass rate61% → 94%
Repeat failures11 → 1
Tokens per run18.4k → 4.2k
Review the results before applying
measured: pending
not deployed
01Surface 01 of 05

Control

Pause, restart, swap the model or push a prompt. The command lands in the integration inside your app.

ACP never routes your LLM traffic. Local state is checked before the next call, so nothing waits on a deploy. Pick a command to send it in the panel.

02Surface 02 of 05

Runs

Open the failing run, read the reason it failed, and get the fix in the same view.

Every run report carries model, latency, tokens, tool calls and the error. Failures are grouped so a repeated cause is obvious instead of buried in a log.

  • Filter to failures, one agent, or one app.
  • Read the error, the explanation of what the agent did, and the steps to fix it.
  • Act from the run: retry it, push a prompt, or send it to the Skills Hub.
03Surface 03 of 05

Token usage

Account for spend per agent, per app, per model, and per day.

Input and output are counted separately, because the fix for a long prompt is not the fix for a long answer. Compare the last 7, 30 or 90 days.

  • Split input against output on every agent and every day.
  • Group by agent, app or model to find the one line that moved.
  • Compare 7, 30 and 90 day windows against the run count behind them.
04Surface 04 of 05

Feedback

Score a response by hand, or let a reviewer model score it. Both teach the agent.

Feedback attaches to the run, not to a spreadsheet. When scores for one agent keep landing low, the same evidence becomes the case for a new skill.

  • Rate any run yourself, with a note the next reviewer can read.
  • Auto-score runs with a reviewer model on resolution, policy and tone.
  • Escalate a pattern of low scores straight into a skill draft.
05Surface 05 of 05

Skills Hub

When an agent keeps failing, it gets taught a skill instead of a patch.

ACP drafts the skill from the failing runs, replays those runs against it in the Skills Testing Lab, and compares the agent before and after. The skill deploys only if the overall result clears the gate.

  • Draft a skill from the failures that prove it is needed.
  • Test it in the lab on replayed runs, before and after, side by side.
  • Deploy on pass only. A skill that does not clear the gate stays a draft.
Observe

Runs, tokens, cost, errors, health.

Every report puts operating context beside the agent that produced it.

Local state

Control without a proxy.

The wrapper checks paused or stopped state before each LLM call.

traffic routeunchanged
agent stateremote
Model override

Change the next run.

The integration resolves the override locally before the request.

beforegpt-4.1-mini
afterclaude-sonnet-4-6
Prompt cache

Fetch, cache, fall back.

Successful fetches are cached; failures fall back to cache, then hardcoded prompt.

currentv12
incomingv13 · cached
Failure path

Your app keeps running.

Heartbeat and report failures are caught so ACP stays outside the critical path.

ACP unavailablecaught
agent runcontinues
Product tour

The screens you operate from.

app.acp / dashboard Simulated data

Dashboard

320 agents across 31 apps

Money

Cost today

$42.18

Tokens today

918.4k

Requests today

1,284

Problems

3

Cost anomaly · 184% above 7-day median

checkout-recovery

ExplainRollbackPause

21 retries in 12 minutes

checkout-recovery

ExplainRollbackPause

No heartbeat for 4 minutes

digest-writer

ExplainRollbackPause

Activity

Recent activity

12scheckout-recoverysuccess1.8k2.4s
41scheckout-recoveryerror—9.1s
1msupport-routersuccess6241.1s
2minvoice-parsersuccess1.1k3.6s
4mlead-enricheridle——

Apps

checkout-suite

14 agents

support-desk

9 agents

billing-ops

6 agents

Live Monitor

Live
Refresh:5s
1 agent needs attention
Search agents…
All
Health (worst)

checkout-recovery

checkout-suite

18gpt-4.1-mini
12s ago184.2k tokens

support-router

support-desk

96gpt-4.1-mini
1m ago42.6k tokens

invoice-parser

billing-ops

91claude-sonnet-4-6
2m ago61.0k tokens

digest-writer

growth-loops

insufficient data
No activity—

Agents

Search agents…
All Apps
All Types
ListHierarchy
AgentAppTypeStatusHealthModelTokensErrorsActions
checkout-recoverycheckout-suiteorchestratorerror18gpt-4.1-mini184.2k6Pause
checkout-retry-workercheckout-suiteworkerrunning64gpt-4.1-mini96.7k—Pause
support-routersupport-deskrouterrunning96gpt-4.1-mini42.6k—Pause
invoice-parserbilling-opstoolidle91claude-sonnet-4-661.0k—Pause
digest-writergrowth-loopsagentpaused—gpt-4.1-mini——Resume
Honest comparison

Understand every run. Act on what you find.

ACP brings run diagnostics, quality scoring, and replay evaluation together with live agent controls. Compare the workflows you need, from investigating a failure to testing a fix and applying an approved change.

Capability ACP Langfuse LangSmith Helicone Braintrust Datadog
Run diagnostics and evaluations Run diagnostics · feedback scoring · replay evaluations Scores & experiments Online & offline evals Tracing Scorers & experiments APM
Retrofit an already-shipped app Generated integration Instrumentation Instrumentation Gateway or SDK Instrumentation Instrumentation
Pause one hardcoded agent remotely Yes Not its job Not its job Not its job Not its job Not its job
Swap model without redeploying Local override App-managed App-managed Via gateway App-managed App-managed
Push prompt without redeploying Versioned fetch + webhook Prompt management Prompt Hub Via gateway / SDK Prompt environments App-managed
Hosted maturity and ecosystem Early Established Established Established Established Enterprise-grade
Swipe table to compare →
Security · fail-open

A control plane you can trust with production.

Remote control is only acceptable if the wire is signed and the failure mode is boring. Both are enforced in the generated integration, not in documentation.

See every security control, our subprocessors, and what we don’t have yet on the Security & Trust page.

Signed commands

HMAC-SHA256 over ts.nonce.body

Compared with a timing-safe check, and verified against the current and previous secret so rotation never drops a command.

Replay defence

300-second window, nonce cache

A stale timestamp is rejected with 403. A nonce that has already been seen is rejected with 403.

Fails open

Reporting errors are caught

Heartbeat and report failures log and return. A prompt fetch failure falls back to cache, then to the hardcoded prompt.

Not in the path

No proxy, no key custody

Your provider keys stay in your app. ACP never sits between your agent and the model.

Objections

The questions a skeptical developer asks.

Answered against what the code actually does today.

No. The integration is a single generated file you paste into the app you already have. Agents stay hardcoded where they are; the wrapper reads their state before each call.

Your agents keep running. Heartbeats and run reports are wrapped in try/catch and only log on failure. Prompt fetches fall back to the local cache, then to the prompt already in your code.

No. There is no gateway and no proxy. A model swap is a local override the wrapper resolves before your own request goes out, so your provider keys never leave your app.

The webhook sets paused state immediately, and the next wrapped run returns before calling the model. The dashboard confirms it on the following heartbeat, within 30 seconds.

No. Deep tracing and evaluation are stronger elsewhere, and the comparison table above says so. Keep them for analysis and use ACP for the intervention they cannot perform.

Zero. The system ran for a year as one operator's production control plane for 320 agents across 31 applications, and the hosted product is at the start of external validation. You would be early.

0 agents operated
0 applications
1 year in daily production
Who built it

Built from the operator’s seat.

The founder designed, built, and operated the system alone for a year as the real control plane for a crypto venture fund. It managed 320 agents across 31 applications in daily production.

That history proves the operating model. ACP currently has zero external connected apps; the public product is at the beginning of that validation.
The whole evaluation

Connect one app. Operate one real agent.

If you can see a heartbeat and pause a live agent without a deploy, you understand the product.

1 Create the appSign up and receive an API key and webhook secret. No card, no sales call.
2 Paste the integrationHand the generated file to your coding agent, or drop it in yourself.
3 Take controlWatch the first heartbeat land, then pause the agent from the dashboard.
your connected apps
checkout-suite14 agents · beat 12s agoconnected
support-desk9 agents · beat 41s agoconnected
your appwaiting for the first heartbeatpending
connect new