This is a working Claude Code routine that triages production Cloudflare Worker errors every 6 hours without anyone watching. Each run pulls the last 6 hours of exceptions from the Workers Observability API, diagnoses each one against the codebase, and then does one of three things: opens a pull request with a tested fix, files a ClickUp task and tags me, or posts an all-clear to Slack. Below are the prompt, the API call, the schedule, and the permission limits I use.
If you have not set up a routine before, start with the complete guide to Claude Code routines. It covers what routines are, the three triggers (schedule, API, GitHub), and the setup flow. This post assumes you know those basics and covers one production use case in detail.
Why error triage was my first routine
I choose automation targets with one question: what do I do on a recurring basis that follows rules I can write down?
Log triage was the first answer. Before the routine, every morning started the same way: open the Cloudflare dashboard, look for exception spikes, check anything unusual against open issues, decide whether it was a real bug or noise, and then start work. That took 20 to 30 minutes a day, and I still missed errors that happened overnight.
The decisions in that process are mostly mechanical:
- Is this error new, or already tracked?
- Is it a code bug, a configuration issue, or a bot probing a route?
- Was it an unhandled crash, or an error the code already caught?
- Is there a safe, testable fix, or does a person need to look at it?
Each question has a rule-based answer, which is the kind of work a scheduled agent can do.
The triage routine, end to end
The routine runs against one of my production apps, a Cloudflare Workers project. Here is the full sequence:
Every 6 hours (fresh cloud session, no prior context): 1. Compute a 360-minute window from `date -u` 2. Run a saved Observability query against the Cloudflare API (production workers only; skip anything -staging / -preview / -dev) 3. Dedupe within the run, then classify each error 4. Read the codebase to diagnose 5a. Isolated, testable code bug -> open ONE PR against `canary` (TDD, never main) 5b. Anything ambiguous -> create a ClickUp task and tag me 6. Post a summary to Slack #claude-logs (every run, even when clean)It has four parts: the query, the prompt, the schedule, and the output contract. The output contract (what the routine is allowed to produce) is the part most people leave undefined, and it matters most.
Step 1: Query the Workers Observability API
The routine calls a saved query in Cloudflare Workers Observability by its id, using curl and three environment variables: $CF_API_TOKEN, $CF_ACCOUNT_ID, and $CF_OBSERVABILITY_QUERY_ID. It does not install wrangler or any other tooling, which keeps each fresh cloud session small and predictable. The routine lists every production worker and skips names ending in -staging, -preview, or -dev.
A routine runs in a fresh cloud session with no access to your local shell, so values in your .zshrc are not available. Add the Cloudflare values to the cloud environment the routine uses. Claude Code's docs note that environment variables are visible to anyone who uses that environment; on Pro and Max plans they recommend storing API keys as API credentials instead. The Default environment also uses Trusted network access, which only allows a fixed list of domains, so check that api.cloudflare.com is reachable and add it under the environment's allowed domains if it is not. If a credential is missing, the prompt tells the routine to stop and post to Slack instead of guessing.
The main call sends the saved query id and a time window to the telemetry endpoint. It filters to records that have an error and groups them by worker, error, and outcome:
curl -sS -X POST \ "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/workers/observability/telemetry/query" \ -H "Authorization: Bearer $CF_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "queryId": "'"$CF_OBSERVABILITY_QUERY_ID"'", "timeframe": { "from": '"$FROM_MS"', "to": '"$TO_MS"' }, "parameters": { "filters": [ { "key": "$metadata.error", "operation": "exists", "type": "string" } ], "filterCombination": "and", "calculations": [{ "operator": "count", "alias": "Count" }], "groupBys": [ { "type": "string", "value": "$workers.scriptName" }, { "type": "string", "value": "$metadata.error" }, { "type": "string", "value": "$workers.outcome" } ], "orderBy": { "value": "Count" }, "limit": 100 } }'The field that matters most is $workers.outcome. A value of exception means an unhandled crash. A value of ok with an error attached means the Worker caught and handled the error. The routine reports both, but treats a handled error as lower severity. To get stack traces, it runs a second query for each affected worker, grouped by the error stack.
If the telemetry query is rejected and the routine cannot fix the request, it falls back to the GraphQL Analytics API (workersInvocationsAdaptive) and notes in the Slack summary that stack traces were unavailable. It does not use the Log Explorer API, which has no Workers dataset.
Nothing here depends on Cloudflare or ClickUp. You can replace the saved Observability query with a Sentry issues feed, a Logpush export in R2, a SQL query against a Supabase logs table, or a Vercel log drain. You can replace ClickUp and Slack with Linear, Jira, GitHub Issues, Discord, or Teams; many of these are available as Claude connectors. The structure stays the same: a saved query produces structured data, the agent classifies it, and the agent proposes changes.
Step 2: Write the routine prompt
The prompt defines what the routine does and what it must not do. This is a trimmed version of mine:
# RoleYou run every 6 hours, unattended, in a fresh cloud session with zero priorcontext. Surface PRODUCTION Cloudflare Worker exceptions from the last 6 hours,diagnose them against the codebase, and either propose a fix via PR or escalateto ClickUp. Read AGENTS.md at the repo root first; it governs architecture, TDD,and the multi-tenancy invariants you must follow for any code change.# Tools & auth- Query the Workers Observability telemetry API with curl, using $CF_API_TOKEN and $CF_ACCOUNT_ID. Do NOT use wrangler or install tooling.- If a token, account id, or query id is missing, or the API returns 401/403: STOP and post the failure to Slack #claude-logs. No fallback, no guessing.- Slack and ClickUp are reached via their MCP connectors.# Scope (hard constraints)- PRODUCTION ONLY. Ignore staging/preview/dev workers entirely.- Window: exceptions from the last 360 minutes, computed from `date -u` at start.- Enumerate every production worker from Cloudflare; do not assume a fixed list.- Never push to main. Never deploy. Never touch infrastructure or secrets.# Procedure1. Compute [from, to] in epoch ms (now-360min .. now).2. POST the saved telemetry query ($CF_OBSERVABILITY_QUERY_ID), filtered to errors, grouped by scriptName + error + outcome. Pull stack traces with a second query per affected worker.3. Dedupe within the run: group by (worker, error type, normalized message, top stack frame); collapse IDs, URLs, and timestamps first. The run is stateless.4. Classify each group: CODE BUG / CONFIG-ENV / TRANSIENT / UNCLEAR.# Act- Open ONE PR against `canary` only for an isolated, well-understood code bug with a safe, testable fix. Follow AGENTS.md TDD (failing test first); run typecheck, lint, and format before opening. Bias toward escalation when unsure.- Otherwise create a ClickUp task assigned to me, titled "[prod-triage] <worker> โ <issue>", with the exception, a stack trace, occurrence count, root-cause analysis, and what needs verifying. Before creating, search for an open "[prod-triage]" task for the same worker and issue; if one exists, comment the new count instead of creating a duplicate.# ReportPost to Slack #claude-logs every run: PRs opened, items needing verification withClickUp links, and transient/no-action notes. If zero exceptions, say so.Most of the prompt is limits. "Bias toward escalation when unsure." "Search for an existing task before creating one." "If a token or query id is missing, stop and post to Slack." A routine that reports noise, or improvises when data is missing, produces reports you stop reading within a week.
Two details from the Claude Code docs affect how you write this prompt. First, all of your connected connectors are included in a new routine by default, and the routine can use every tool in them, including writes, without asking. Remove any connector the prompt does not need. Second, commits, pull requests, and connector actions appear under your own GitHub and Slack identities, so teammates will see them as coming from you.
Step 3: Set the schedule
The routine runs every 6 hours: 0 */6 * * * in cron syntax. At the start of each run it computes a 360-minute lookback window from date -u, so the query window always matches the interval between runs.
The routine form only offers presets (hourly, daily, weekdays, weekly). To use a custom interval such as every 6 hours, pick the closest preset, then run /schedule update in the Claude Code CLI and set the cron expression. The minimum interval is one hour.
Choose the interval based on how long a problem can go unnoticed before it hurts users. Six hours fits my traffic: errors do not sit overnight, and the number of runs stays low. On a higher-traffic app I would run it hourly; on a side project, daily. Routines also have a daily cap on runs per account, so check your remaining runs at claude.ai/code/routines before tightening the schedule across several routines.
Step 4: Define the output contract
The output contract decides whether the routine helps or causes damage. Mine allows three outcomes:
- Isolated, testable code bug: open one PR against
canarywith a failing test first, after running typecheck, lint, and format. I review and merge it. - Anything ambiguous or risky: create a ClickUp task with the evidence and tag me. The prompt tells the agent to escalate whenever it is unsure.
- Clean run: post one line to Slack and stop.
The routine cannot push to main and cannot deploy. The most it can do is open a PR against canary, a branch that does not deploy on its own, and a person reviews every merge. Claude Code enforces part of this by default: a routine pushes its work to claude/-prefixed branches and rejects pushes to protected branches. Everything else the routine does is write notes for me to read. Because of that limit, I am comfortable pointing it at production logs.
Three triage rules to copy
Setting up the routine is straightforward. These three rules decide whether its reports are useful.
Deduplicate within each run. Runs are stateless, so the routine cannot remember what it reported last time. It groups errors by worker, error type, normalized message, and top stack frame, and it strips IDs, URLs, and timestamps from messages before grouping. Without this step, one ongoing incident appears as 50 separate rows.
Separate handled errors from crashes. In Cloudflare, outcome=exception is an unhandled crash, and outcome=ok with an error means your code caught it. Putting that distinction in the prompt stops the agent from escalating errors your code already handles.
Escalate by default. The agent opens a PR only for an isolated bug with a clear, testable fix. Everything else becomes a ClickUp task. A wrong PR costs more review time than a task that asks a question, so the prompt favors escalation.
What the routine is allowed to do
I set routine permissions the way I would for a junior engineer in their first week: plenty of access to investigate and propose, none to change production.
Give a routine write access for:
- Opening PRs against a branch that does not deploy (a person reviews the merge)
- Creating issues or tasks
- Adding labels and comments
- Posting run summaries to a Slack channel
Do not give it write access for:
- Pushing to
mainor any branch that deploys - Closing issues or resolving tasks (it does not have the full context)
- Messaging customers or notifying the whole team
- Changing infrastructure, billing, secrets, or auth configuration
If unsure: let the routine propose, not act. A PR or a task is a proposal; a merge or a deploy is an action. Keep actions with a person, and make the routine stop and post to Slack as soon as it lacks the access or data it needs.
Results so far, and the risks I am tracking
The routine is still new, so I can report what it has found, not long-term metrics.
It has been most useful for problems I would not have looked for. It found a misconfiguration I did not know about, and it regularly reports bad bots: requests that probe routes and generate errors. That pattern does not stand out on the dashboard, but it is obvious when an agent reads every production worker's exceptions on a schedule.
Three risks I expect to hit but have not yet:
- Severity inflation. Agents tend to rate every error as important. The outcome-based classification and the escalate-by-default rule are meant to limit this, and I will tighten them if the routine over-reports.
- Treating external outages as bugs. A third-party outage can look like a bug in your code. Deduplication and the "classify as transient, do not act" rule help, but I trust this part of the routine least.
- Cost growth. Routines use your subscription's usage like an interactive session. Four runs a day is cheap; hourly runs across more workers is not. Measure usage before you increase the frequency.
Other routines with the same structure
The same structure (a saved input, rules written in the prompt, output limited to proposals) fits other recurring work:
- Dependency triage: weekly, read the changelog and lockfile diff for outdated packages, open a PR for safe patch bumps, and file a task for breaking changes.
- Daily SEO brief: every morning, pull ranking and traffic changes, summarize what moved, and flag pages that dropped.
- PR hygiene: a few times a day, find PRs that are stale, missing a description, or failing CI, and leave a comment.
- Docs drift: weekly, compare recent code changes against the docs and open tasks where they disagree.
Routine vs cron script vs interactive agent
A routine is not always the right tool:
| Tool | Best for | Avoid when |
|---|---|---|
| Cron script | Deterministic work with no judgment (backups, syncs) | The task needs reasoning about context |
| Claude Code routine | Recurring work that needs judgment and can propose changes | The task is one-off, or needs your input mid-task |
| Interactive agent | One-off or exploratory work where you steer | The work is recurring and you keep redoing it |
If a plain script can do the job, use the script. Use a routine when the task requires reading a situation and deciding what to do. The routines guide extends this comparison to subagents and Managed Agents.
Quick Recommendation
Claude Code routines are best for:
- Recurring work that follows rules you can write down (triage, audits, briefs)
- Tasks that should happen on schedule whether or not you remember them
- Cases where a proposal (PR, task, Slack summary) is more useful than a raw alert
Skip routines if:
- The task is one-off or needs you to steer it while it runs
- A deterministic cron script already handles it
- You cannot define a strict output contract (a routine without one produces noise)
My pick: start with one read-only routine that only files tasks and posts to Slack. Run it for a week, and grant PR access only after its classifications have been correct. Add write access in steps, as you would for a new hire.
Frequently Asked Questions
What are Claude Code routines?
How are Claude Code routines different from a cron job?
How do I pull Cloudflare logs into a routine?
Is it safe to give a scheduled agent access to production logs?
How often should a triage routine run?
What stops the routine from sending false positives?
Next steps
For the foundations, read our Claude Code best practices guide on how we structure AGENTS.md rules, MCP servers, and subagents, then how to build a SaaS with Claude Code for the interactive version of this workflow. If you are choosing where to host the app a routine like this monitors, see best hosting for Next.js.
MakerKit ships with AGENTS.md rules for AI agents and an MCP server, which give a routine the project context it needs to diagnose errors and write fixes that follow the codebase's conventions.
One last note. Nothing stops you from removing the canary limit and letting the routine push straight to production. I would not, and the rest of this post explains why. The routine will operate within whatever permissions you give it, so choose them deliberately.