The product that improves itself.

Drums watches how people use your product in production: where it fails, where they get stuck, where it could do better. It writes down what it expects before a change ships, makes the change with the coding agent you already use, and measures whether it helped. Every change is proposed to you first. The record is waiting with your coffee.

Get early access
Trusted by high-horsepower founding teams at
twentyfour26 Layman Flocta and more

Bring your own agent. Drums runs the loop.

Claude Code, Codex, Gemini CLI, Cursor, Amp and OpenCode write the change. Drums decides what to change, whether it helped, and whether it ships.

Why Drums stays agent-neutral

Claude Code Codex Gemini CLI Cursor Amp OpenCode
The thesis

Don't declare it better.
Measure that it helped.

Drums starts by only watching. Nothing changes until you say so.

Every claim says how it is known. Never one green check.

A metric that moved is not proof the change caused it. Drums says how sure it can be, and a full rollout never proves cause.

Observe first
0
changes before you say so
watching every deploychanging nothing yet
How it knows
5states
verifiedobservedinferredapprovedunresolved
verified · Drums ran it and watched it pass observed · telemetry says so inferred · a model concluded it; goes to a person approved · a named person signed off unresolved · Drums does not know, and says so
Undo
30days
every change reversible with one command
undo · printed with the resultdrums undo --since 1h reverses itdrums stop pauses everything

Measured, not declared

Every change carries the plan it shipped under: the outcome it expects, the baseline, the window. It is measured once, when that window closes, and never re-cut to look better. Drums checks back at 7, 30, and 90 days, because a change that looked good at close can quietly regress later.

Says how it knows, not green checks

Verified, observed, inferred, approved, unresolved. Every claim carries one of the five, and only verified or observed can ship alone.

Git, not another database

Changes land as real commits. The evidence, meaning what was seen, what was changed, how it was verified and who approved, attaches as a git note. git log is the audit trail.

Earned autonomy, not a settings page

Each failure class climbs from observe to act-alone on its own record, and drops a rung automatically after rollbacks.

Approval, not blanket autonomy

Customer data, billing, permissions, and infrastructure always wait for a named person.

Agent-neutral, no lock-in

Claude Code and Codex write the repair today, OpenCode next. Switching agents changes nothing about the record.

How the loop works

01

Reproduce

A failing request is captured from your running app and tied to the deploy that shipped just before it. Then Drums makes the failure happen again, so the link is a fact, not a guess.

Detected from telemetry

An observation arrives over an event webhook, with the evidence attached. Drums reads your PostHog to measure outcomes today; Sentry and OpenTelemetry are next.

Attributed to a deploy

The failing stack trace is matched against the deploy that shipped just before it and the files it changed.

Rebuilt at that revision

An isolated copy of the app is rebuilt at the exact commit the failure first appeared on.

Replayed, not guessed

The captured request runs again. If it doesn't fail, Drums says so instead of writing a patch.

Intent carried forward

The PR body, the linked ticket, or the agent prompt behind the change is kept as what it was supposed to do.

The Drums review surface: a queue of incidents on the left; an open incident with its evidence, captured request, and a verified fix waiting on one approval.
The review surface, from the current build: a seeded lab workspace, the same UI a design partner signs into.
Reproduction is what turns an attribution into a fact. The loop, end to end
02

Repair

Your own coding agent writes the change. Drums hands it the observation, the hypothesis, and the acceptance criteria, then checks the result against real behaviour, replaying the original failure where there is one. Never against the agent's report.

Rolling out with design partners now; the docs say exactly what ships today.

Repair · checkout · 8f32a1verified
CheckResult
original failing request✓ verified · returns 200
failure signature✓ verified · match at 8f32a1
build at the repair✓ verified · boots clean
Every claim is tied to the exact revision, environment, and actor.
03

Release

The change goes to a small share of traffic first and is measured against the observation that prompted it: the failure gone, the number moved. Then it is promoted or rolled back, and shipped changes stay reversible with one command for 30 days.

In design; the docs say exactly what ships today.

Repair verified Canary Compare Promote · recovered Revert · regressed

Every step, decision, and revert lands in the record.

Canary · 10% of trafficholding · 90s
errors on the failing route3.1% → 0.0%
p95 latency312ms → 308ms
Promotion waits until the original failure is gone and nothing new appears.
Promoteheld
new error signaturesnone
adjacent routesunchanged
full rollout△ held until every check holds
Promotion halts unless every check holds.
Revertreversible
undo commandprinted with the result
reversible window30 days
drums stoppauses everywhere
Reversibility is stated at the moment of the action, not in the documentation.
Approvala named person
customer data, billing, permissions, infra△ always ask
approvalsigned by a named person
agent self-approval✕ never
The sensitive paths never ship without a person.

Deployed by your systems.

Drums drives the deploy platform you already run. It never becomes your deploy target, and a held promotion cannot be silently overridden.

Stop and undo
Safety

Built so trust survives the first mistake.

Every claim says how it is known. Autonomy is earned per failure class. One command stops everything, and approval is required where a mistake would cost the most.

Every claim says how it knows

Verified, observed, inferred, approved, or unresolved. Never one green check.

Boundaries in plain language

Proposed from your reverts, code owners, and migration history. Corrected in a sentence, not a config file.

Git is the record

Changes land as real commits. The evidence, meaning what was seen, what was changed, how it was verified and who approved, attaches as a git note that travels with the repository.

Approval where it counts

Customer data, billing, permissions, and infrastructure always wait for a named person.

Verified against the failure

The original failing request, replayed against the repair until the failure is gone. Not the agent's summary.

Reproduction inside verification

An isolated copy of that exact revision, the captured request replayed. The guess becomes a fact before anything is written.

Canary before promotion

A small share of traffic first. Promotion waits until the original failure is gone and nothing new appears.

One command to stop

drums stop pauses every environment; drums undo --since reverses what shipped inside a window.

Shadow mode

Repairs generated and verified in isolation, never shipped, so you can compare them against what your team actually did.

Autonomy earned per class

Observe → shadow → propose → act alone, promoted on track record and demoted automatically after rollbacks.

One binary

A CLI today; a git remote and an MCP tool inside Claude Code and Codex are in design. No new place to go.

Built on your systems

A webhook in today; Sentry, PostHog, and OpenTelemetry adapters next. Your existing deploy platform out. Drums never becomes your deploy target.

Pricing

Start free. Pay when it works for you.

The whole loop short of acting alone is free on your own machine. Paid plans add the hosted record, your team, and rollouts that measure themselves. Drums is invite-only right now: every paid seat starts with a call.

Developer

The loop on your own machine

Freeno credit card
  • Get started with:
  • Unlimited repositories
  • Observe, reproduce, propose
  • Your own coding agent writes the change
  • Every claim labeled with how it is known
  • The record stays local
Install the CLI

Enterprise

For organizations with boundaries

Custom
  • Everything in Pro, plus:
  • Private deployment, your infrastructure
  • Identity-bound approvals and SSO
  • Signed audit export
  • Volume pricing across applications
Contact us

The engineer is on the loop, not in it.

Drums proposes by default; you approve. Every shipped change stays one command away from undone.

FAQ

The self-improving product, answered plainly

What Drums does, what it refuses to do on its own, and where it stops.

What is a self-improving product?
A product that watches how people use it in production, finds problems and opportunities its own code can address, makes the change, and measures whether it actually helped, without a person having to notice first. Drums runs that as a loop: observe, understand, form a hypothesis, change with your own coding agent, verify, roll out, measure, learn.
How does Drums fix production bugs automatically?
It replays the failing request against a container built at the commit that introduced the failure. If it does not fail there, Drums says so and stops rather than guessing. If it does, the reproduction becomes the acceptance test the repair has to pass, so a fix is only called verified when the original failure has actually stopped.
Does Drums deploy to production without approval?
No, not by default. Every repair stops at a pull request and waits. Shipping on its own is earned per failure class: five consecutive clean ships, a human running drums authority promote, and automatic demotion the first time something goes wrong. There is no flag that skips the ladder.
Which coding agents does it work with?
Claude Code and Codex today, OpenCode next. Drums drives the CLI already authenticated on your machine. It never resells model access and never asks for an API key in local mode. Deciding what needs repairing and whether the repair worked is the product; writing the patch is not.
How is this different from Sentry or Datadog?
They are inputs, not competitors. They tell you what is happening. Today Drums detects through its own capture in your app, repairs what it can reproduce, and proposes the change for your approval. It reads the PostHog you already run to measure whether a shipped change actually helped. Readers for Sentry and OpenTelemetry are in development. It has no graphs, alert rules, or on-call schedule of its own.
What does it cost?
Developer is free forever: the whole loop short of acting alone, on your own machine, on as many repositories as you like, with no account. Pro is priced per production application, and the number comes from a call — every paid seat starts with one. It adds the hosted record, team approvals, and measured rollouts. Enterprise is custom: private deployment, identity-bound approvals, and a signed audit export.
Founders

What founders say about Drums

“Self Healing Software” have been testing drums for a while now, it has accelerated our development workflow at twentyfour26 by a huge margin
Sam Mathew Cofounder & CTO, twentyfour26