ยทUpdated

Claude Code Routine Example: Automated Production Error Triage (2026)

A real Claude Code routine that triages production Cloudflare Worker errors every 6 hours: the full prompt, the API call, the schedule, and the permission limits that keep it safe.

This is a working Claude Code routine that triages production Cloudflare Worker errors every 6 hours without anyone watching. Each run pulls the last 6 hours of exceptions from the Workers Observability API, diagnoses each one against the codebase, and then does one of three things: opens a pull request with a tested fix, files a ClickUp task and tags me, or posts an all-clear to Slack. Below are the prompt, the API call, the schedule, and the permission limits I use.

If you have not set up a routine before, start with the complete guide to Claude Code routines. It covers what routines are, the three triggers (schedule, API, GitHub), and the setup flow. This post assumes you know those basics and covers one production use case in detail.

Why error triage was my first routine

I choose automation targets with one question: what do I do on a recurring basis that follows rules I can write down?

Log triage was the first answer. Before the routine, every morning started the same way: open the Cloudflare dashboard, look for exception spikes, check anything unusual against open issues, decide whether it was a real bug or noise, and then start work. That took 20 to 30 minutes a day, and I still missed errors that happened overnight.

The decisions in that process are mostly mechanical:

  • Is this error new, or already tracked?
  • Is it a code bug, a configuration issue, or a bot probing a route?
  • Was it an unhandled crash, or an error the code already caught?
  • Is there a safe, testable fix, or does a person need to look at it?

Each question has a rule-based answer, which is the kind of work a scheduled agent can do.

The triage routine, end to end

The routine runs against one of my production apps, a Cloudflare Workers project. Here is the full sequence:

Every 6 hours (fresh cloud session, no prior context):
1. Compute a 360-minute window from `date -u`
2. Run a saved Observability query against the Cloudflare API
(production workers only; skip anything -staging / -preview / -dev)
3. Dedupe within the run, then classify each error
4. Read the codebase to diagnose
5a. Isolated, testable code bug -> open ONE PR against `canary` (TDD, never main)
5b. Anything ambiguous -> create a ClickUp task and tag me
6. Post a summary to Slack #claude-logs (every run, even when clean)

It has four parts: the query, the prompt, the schedule, and the output contract. The output contract (what the routine is allowed to produce) is the part most people leave undefined, and it matters most.

Step 1: Query the Workers Observability API

The routine calls a saved query in Cloudflare Workers Observability by its id, using curl and three environment variables: $CF_API_TOKEN, $CF_ACCOUNT_ID, and $CF_OBSERVABILITY_QUERY_ID. It does not install wrangler or any other tooling, which keeps each fresh cloud session small and predictable. The routine lists every production worker and skips names ending in -staging, -preview, or -dev.

The main call sends the saved query id and a time window to the telemetry endpoint. It filters to records that have an error and groups them by worker, error, and outcome:

curl -sS -X POST \
"https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/workers/observability/telemetry/query" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"queryId": "'"$CF_OBSERVABILITY_QUERY_ID"'",
"timeframe": { "from": '"$FROM_MS"', "to": '"$TO_MS"' },
"parameters": {
"filters": [
{ "key": "$metadata.error", "operation": "exists", "type": "string" }
],
"filterCombination": "and",
"calculations": [{ "operator": "count", "alias": "Count" }],
"groupBys": [
{ "type": "string", "value": "$workers.scriptName" },
{ "type": "string", "value": "$metadata.error" },
{ "type": "string", "value": "$workers.outcome" }
],
"orderBy": { "value": "Count" },
"limit": 100
}
}'

The field that matters most is $workers.outcome. A value of exception means an unhandled crash. A value of ok with an error attached means the Worker caught and handled the error. The routine reports both, but treats a handled error as lower severity. To get stack traces, it runs a second query for each affected worker, grouped by the error stack.

If the telemetry query is rejected and the routine cannot fix the request, it falls back to the GraphQL Analytics API (workersInvocationsAdaptive) and notes in the Slack summary that stack traces were unavailable. It does not use the Log Explorer API, which has no Workers dataset.

Step 2: Write the routine prompt

The prompt defines what the routine does and what it must not do. This is a trimmed version of mine:

# Role
You run every 6 hours, unattended, in a fresh cloud session with zero prior
context. Surface PRODUCTION Cloudflare Worker exceptions from the last 6 hours,
diagnose them against the codebase, and either propose a fix via PR or escalate
to ClickUp. Read AGENTS.md at the repo root first; it governs architecture, TDD,
and the multi-tenancy invariants you must follow for any code change.
# Tools & auth
- Query the Workers Observability telemetry API with curl, using $CF_API_TOKEN
and $CF_ACCOUNT_ID. Do NOT use wrangler or install tooling.
- If a token, account id, or query id is missing, or the API returns 401/403:
STOP and post the failure to Slack #claude-logs. No fallback, no guessing.
- Slack and ClickUp are reached via their MCP connectors.
# Scope (hard constraints)
- PRODUCTION ONLY. Ignore staging/preview/dev workers entirely.
- Window: exceptions from the last 360 minutes, computed from `date -u` at start.
- Enumerate every production worker from Cloudflare; do not assume a fixed list.
- Never push to main. Never deploy. Never touch infrastructure or secrets.
# Procedure
1. Compute [from, to] in epoch ms (now-360min .. now).
2. POST the saved telemetry query ($CF_OBSERVABILITY_QUERY_ID), filtered to
errors, grouped by scriptName + error + outcome. Pull stack traces with a
second query per affected worker.
3. Dedupe within the run: group by (worker, error type, normalized message, top
stack frame); collapse IDs, URLs, and timestamps first. The run is stateless.
4. Classify each group: CODE BUG / CONFIG-ENV / TRANSIENT / UNCLEAR.
# Act
- Open ONE PR against `canary` only for an isolated, well-understood code bug
with a safe, testable fix. Follow AGENTS.md TDD (failing test first); run
typecheck, lint, and format before opening. Bias toward escalation when unsure.
- Otherwise create a ClickUp task assigned to me, titled "[prod-triage] <worker>
โ€” <issue>", with the exception, a stack trace, occurrence count, root-cause
analysis, and what needs verifying. Before creating, search for an open
"[prod-triage]" task for the same worker and issue; if one exists, comment the
new count instead of creating a duplicate.
# Report
Post to Slack #claude-logs every run: PRs opened, items needing verification with
ClickUp links, and transient/no-action notes. If zero exceptions, say so.

Most of the prompt is limits. "Bias toward escalation when unsure." "Search for an existing task before creating one." "If a token or query id is missing, stop and post to Slack." A routine that reports noise, or improvises when data is missing, produces reports you stop reading within a week.

Two details from the Claude Code docs affect how you write this prompt. First, all of your connected connectors are included in a new routine by default, and the routine can use every tool in them, including writes, without asking. Remove any connector the prompt does not need. Second, commits, pull requests, and connector actions appear under your own GitHub and Slack identities, so teammates will see them as coming from you.

Step 3: Set the schedule

The routine runs every 6 hours: 0 */6 * * * in cron syntax. At the start of each run it computes a 360-minute lookback window from date -u, so the query window always matches the interval between runs.

The routine form only offers presets (hourly, daily, weekdays, weekly). To use a custom interval such as every 6 hours, pick the closest preset, then run /schedule update in the Claude Code CLI and set the cron expression. The minimum interval is one hour.

Choose the interval based on how long a problem can go unnoticed before it hurts users. Six hours fits my traffic: errors do not sit overnight, and the number of runs stays low. On a higher-traffic app I would run it hourly; on a side project, daily. Routines also have a daily cap on runs per account, so check your remaining runs at claude.ai/code/routines before tightening the schedule across several routines.

Step 4: Define the output contract

The output contract decides whether the routine helps or causes damage. Mine allows three outcomes:

  • Isolated, testable code bug: open one PR against canary with a failing test first, after running typecheck, lint, and format. I review and merge it.
  • Anything ambiguous or risky: create a ClickUp task with the evidence and tag me. The prompt tells the agent to escalate whenever it is unsure.
  • Clean run: post one line to Slack and stop.

The routine cannot push to main and cannot deploy. The most it can do is open a PR against canary, a branch that does not deploy on its own, and a person reviews every merge. Claude Code enforces part of this by default: a routine pushes its work to claude/-prefixed branches and rejects pushes to protected branches. Everything else the routine does is write notes for me to read. Because of that limit, I am comfortable pointing it at production logs.

Three triage rules to copy

Setting up the routine is straightforward. These three rules decide whether its reports are useful.

Deduplicate within each run. Runs are stateless, so the routine cannot remember what it reported last time. It groups errors by worker, error type, normalized message, and top stack frame, and it strips IDs, URLs, and timestamps from messages before grouping. Without this step, one ongoing incident appears as 50 separate rows.

Separate handled errors from crashes. In Cloudflare, outcome=exception is an unhandled crash, and outcome=ok with an error means your code caught it. Putting that distinction in the prompt stops the agent from escalating errors your code already handles.

Escalate by default. The agent opens a PR only for an isolated bug with a clear, testable fix. Everything else becomes a ClickUp task. A wrong PR costs more review time than a task that asks a question, so the prompt favors escalation.

What the routine is allowed to do

I set routine permissions the way I would for a junior engineer in their first week: plenty of access to investigate and propose, none to change production.

Give a routine write access for:

  • Opening PRs against a branch that does not deploy (a person reviews the merge)
  • Creating issues or tasks
  • Adding labels and comments
  • Posting run summaries to a Slack channel

Do not give it write access for:

  • Pushing to main or any branch that deploys
  • Closing issues or resolving tasks (it does not have the full context)
  • Messaging customers or notifying the whole team
  • Changing infrastructure, billing, secrets, or auth configuration

If unsure: let the routine propose, not act. A PR or a task is a proposal; a merge or a deploy is an action. Keep actions with a person, and make the routine stop and post to Slack as soon as it lacks the access or data it needs.

Results so far, and the risks I am tracking

The routine is still new, so I can report what it has found, not long-term metrics.

It has been most useful for problems I would not have looked for. It found a misconfiguration I did not know about, and it regularly reports bad bots: requests that probe routes and generate errors. That pattern does not stand out on the dashboard, but it is obvious when an agent reads every production worker's exceptions on a schedule.

Three risks I expect to hit but have not yet:

  • Severity inflation. Agents tend to rate every error as important. The outcome-based classification and the escalate-by-default rule are meant to limit this, and I will tighten them if the routine over-reports.
  • Treating external outages as bugs. A third-party outage can look like a bug in your code. Deduplication and the "classify as transient, do not act" rule help, but I trust this part of the routine least.
  • Cost growth. Routines use your subscription's usage like an interactive session. Four runs a day is cheap; hourly runs across more workers is not. Measure usage before you increase the frequency.

Other routines with the same structure

The same structure (a saved input, rules written in the prompt, output limited to proposals) fits other recurring work:

  • Dependency triage: weekly, read the changelog and lockfile diff for outdated packages, open a PR for safe patch bumps, and file a task for breaking changes.
  • Daily SEO brief: every morning, pull ranking and traffic changes, summarize what moved, and flag pages that dropped.
  • PR hygiene: a few times a day, find PRs that are stale, missing a description, or failing CI, and leave a comment.
  • Docs drift: weekly, compare recent code changes against the docs and open tasks where they disagree.

Routine vs cron script vs interactive agent

A routine is not always the right tool:

ToolBest forAvoid when
Cron scriptDeterministic work with no judgment (backups, syncs)The task needs reasoning about context
Claude Code routineRecurring work that needs judgment and can propose changesThe task is one-off, or needs your input mid-task
Interactive agentOne-off or exploratory work where you steerThe work is recurring and you keep redoing it

If a plain script can do the job, use the script. Use a routine when the task requires reading a situation and deciding what to do. The routines guide extends this comparison to subagents and Managed Agents.

Quick Recommendation

Claude Code routines are best for:

  • Recurring work that follows rules you can write down (triage, audits, briefs)
  • Tasks that should happen on schedule whether or not you remember them
  • Cases where a proposal (PR, task, Slack summary) is more useful than a raw alert

Skip routines if:

  • The task is one-off or needs you to steer it while it runs
  • A deterministic cron script already handles it
  • You cannot define a strict output contract (a routine without one produces noise)

My pick: start with one read-only routine that only files tasks and posts to Slack. Run it for a week, and grant PR access only after its classifications have been correct. Add write access in steps, as you would for a new hire.

Frequently Asked Questions

What are Claude Code routines?
Claude Code routines are saved Claude Code configurations (a prompt, one or more repositories, and a set of connectors) that run automatically on Anthropic-managed cloud infrastructure. A routine starts from a schedule, an API call, or a GitHub event, runs without a person watching, and each run is a fresh session. Use one for recurring, rule-based work like log triage, dependency checks, or daily briefs.
How are Claude Code routines different from a cron job?
A cron job runs fixed, deterministic steps. A Claude Code routine runs an agent that reasons about context: it can read your logs and codebase, decide whether an error is a real bug or bot noise, and propose a fix as a pull request. Use a script when no judgment is needed and a routine when the task requires reading a situation and deciding.
How do I pull Cloudflare logs into a routine?
I query the Workers Observability telemetry API with curl, authenticating with environment variables and calling a saved query by id so the data shape stays consistent every run. The GraphQL Analytics API is a fallback when the telemetry query is rejected. Avoid the Log Explorer API, which has no Workers dataset, and avoid wrangler so the cloud session stays minimal. Make sure api.cloudflare.com is allowed in the cloud environment's network settings.
Is it safe to give a scheduled agent access to production logs?
Yes, if you limit what it can do. My routine reads logs and opens PRs against a non-deploying branch (canary), but it cannot push to main, deploy, close tasks, message customers, or change infrastructure. It stops and posts to Slack if its credentials or query id are missing instead of guessing. Remove connectors the routine does not need, because it can use every included connector's tools without asking.
How often should a triage routine run?
Match the interval to how long a problem can go unnoticed before it hurts users. I run mine every 6 hours with a matching 360-minute lookback window. Higher-traffic apps can run hourly, which is the minimum interval; side projects are fine once a day. Custom intervals like every 6 hours are set with /schedule update in the CLI.
What stops the routine from sending false positives?
Three rules: deduplicate within each run by worker, error type, normalized message, and top stack frame; separate handled errors (outcome ok) from unhandled crashes (outcome exception) when setting severity; and create a ClickUp task instead of a PR whenever a fix is not clear and testable. A routine that reports noise on a clean run gets ignored within a week.

Next steps

For the foundations, read our Claude Code best practices guide on how we structure AGENTS.md rules, MCP servers, and subagents, then how to build a SaaS with Claude Code for the interactive version of this workflow. If you are choosing where to host the app a routine like this monitors, see best hosting for Next.js.

MakerKit ships with AGENTS.md rules for AI agents and an MCP server, which give a routine the project context it needs to diagnose errors and write fixes that follow the codebase's conventions.

One last note. Nothing stops you from removing the canary limit and letting the routine push straight to production. I would not, and the rest of this post explains why. The routine will operate within whatever permissions you give it, so choose them deliberately.