Skip to content
Awsaf Alam
GitHubLinkedIn

Cloudflare Workers + D1 as a backend for AI agents

— Cloudflare, Serverless, AI Agents — 9 min read

Most of what an AI agent does is wait. It waits on the model, waits on a search API, waits on another model call. The CPU is idle most of the time. That makes agents a good fit for a platform that bills mainly for CPU time and runs close to users, and a bad fit for a platform that kills any request longer than a few seconds.

Cloudflare Workers sits in an interesting spot between those two. This post covers how I put an agent backend together from Workers, D1 and Workflows, the limits that shape the design, and the cases where I'd pick a regular server and Postgres instead.

All limits quoted here come from Cloudflare's docs in July 2026. They change, so check the Workers limits, D1 limits and Workflows limits pages before you rely on a number.

The pieces

Four rows. Worker: one request that finishes while the client waits. Workflow, highlighted: many steps over minutes to days, with retries and sleeps. Durable Object: one live session, such as chat state or WebSockets. D1: records you query across runs and users.
  • Workers handle HTTP. On the paid plan, CPU time per request defaults to 30 seconds and can be raised to 5 minutes. Wall-clock time has no fixed limit while the client stays connected, and time spent waiting on fetch doesn't count as CPU.
  • D1 is SQLite behind a binding. Up to 10 GB per database on the paid plan, with a cap on queries per Worker invocation (1,000 on paid).
  • Workflows run multi-step jobs durably. Each step's result is saved, failed steps retry on their own, and an instance can sleep for hours or days.
  • Durable Objects give you a single-threaded object with its own storage, addressed by ID. They're the right home for one live conversation, especially over WebSockets. Cloudflare's Agents SDK is built on them.

For a "submit a task, get an answer later" agent, I start with the first three.

The shape

Architecture diagram. The client posts a task to a Worker and later polls it. The Worker inserts a run row into a D1 runs table, then creates a Workflow instance with the run ID. Inside the Workflow instance, which is durable and retried, steps run in order: model call 1, tool call 1, model call 2, and finally a save step. Steps call out to LLM and tool APIs and write status back to D1.
  1. The client sends POST /runs. The Worker validates the input, inserts a runs row with status queued, starts a Workflow instance whose ID is the run ID, and returns 202 right away.
  2. The Workflow runs the agent loop. Each model call and each tool call is its own step.
  3. The final step writes the answer and status to D1.
  4. The client polls GET /runs/:id, or you push a notification when the run finishes.

The Worker never holds a request open for the length of an agent run. That alone gets rid of most timeout problems.

The schema

D1 uses SQLite syntax. Put this in a migration file and apply it with wrangler d1 migrations apply:

sql
CREATE TABLE IF NOT EXISTS runs (
  id         TEXT PRIMARY KEY,
  task       TEXT NOT NULL,
  status     TEXT NOT NULL,   -- queued | running | done | stopped | failed
  answer     TEXT,
  error      TEXT,
  created_at INTEGER NOT NULL
);

The Worker

ts
export default {
  async fetch(req: Request, env: Env): Promise<Response> {
    const url = new URL(req.url)
 
    if (req.method === 'POST' && url.pathname === '/runs') {
      const body = await req.json<{ task?: unknown }>().catch(() => ({ task: undefined }))
      if (typeof body.task !== 'string' || body.task.length === 0 || body.task.length > 4000) {
        return Response.json({ error: 'task must be a non-empty string' }, { status: 400 })
      }
      const runId = crypto.randomUUID()
      await env.DB.prepare('INSERT INTO runs (id, task, status, created_at) VALUES (?, ?, ?, ?)')
        .bind(runId, body.task, 'queued', Date.now())
        .run()
      await env.AGENT_RUN.create({ id: runId, params: { task: body.task } })
      return Response.json({ runId }, { status: 202 })
    }
 
    const match = url.pathname.match(/^\/runs\/([\w-]+)$/)
    if (req.method === 'GET' && match) {
      const row = await env.DB.prepare('SELECT id, status, answer, error FROM runs WHERE id = ?')
        .bind(match[1])
        .first()
      return row ? Response.json(row) : Response.json({ error: 'not found' }, { status: 404 })
    }
 
    return Response.json({ error: 'not found' }, { status: 404 })
  },
}

Always use prepare(...).bind(...). Never build SQL out of strings that came from a user or a model. In a real deployment this route also needs authentication, which I've left out to keep the example short.

The Workflow

ts
import { WorkflowEntrypoint, type WorkflowEvent, type WorkflowStep, type WorkflowStepConfig } from 'cloudflare:workers'
 
type RunParams = { task: string }
 
interface Env {
  DB: D1Database
  AGENT_RUN: Workflow<RunParams>
  ANTHROPIC_API_KEY: string
}
 
// Your code: callModel(apiKey, messages) => Promise<Turn>
//            runTool(env, name, input)   => Promise<string>
// Content is plain text here to keep the sketch short.
type Msg = { role: 'user' | 'assistant'; content: string }
type Turn = { text: string; toolCall: { name: string; input: string } | null }
 
const MAX_TURNS = 12
const RETRY = {
  retries: { limit: 3, delay: '10 seconds', backoff: 'exponential' },
  timeout: '2 minutes',
} satisfies WorkflowStepConfig
 
export class AgentRun extends WorkflowEntrypoint<Env, RunParams> {
  async run(event: WorkflowEvent<RunParams>, step: WorkflowStep) {
    const runId = event.instanceId
    const db = this.env.DB
    const setStatus = (status: string, answer: string | null, error: string | null) =>
      db.prepare('UPDATE runs SET status = ?, answer = ?, error = ? WHERE id = ?')
        .bind(status, answer, error, runId)
        .run()
 
    await step.do('mark running', async () => {
      await setStatus('running', null, null)
    })
 
    const messages: Msg[] = [{ role: 'user', content: event.payload.task }]
    let answer: string | null = null
    try {
      for (let turn = 1; turn <= MAX_TURNS; turn++) {
        const reply = await step.do(`model ${turn}`, RETRY, () =>
          callModel(this.env.ANTHROPIC_API_KEY, messages),
        )
        messages.push({ role: 'assistant', content: reply.text })
        if (!reply.toolCall) {
          answer = reply.text
          break
        }
        const call = reply.toolCall
        const result = await step.do(`tool ${turn}: ${call.name}`, RETRY, () =>
          runTool(this.env, call.name, call.input),
        )
        messages.push({ role: 'user', content: result })
      }
    } catch (err) {
      // A step ran out of retries. Record it, then let the instance fail.
      await step.do('save failure', async () => {
        await setStatus('failed', null, String(err))
      })
      throw err
    }
 
    await step.do('save result', async () => {
      if (answer === null) await setStatus('stopped', null, 'max_turns')
      else await setStatus('done', answer, null)
    })
  }
}

And the bindings in wrangler.jsonc:

jsonc
{
  "name": "agent-backend",
  "main": "src/index.ts",
  "compatibility_date": "2026-07-01",
  "d1_databases": [{ "binding": "DB", "database_name": "agent-db", "database_id": "<your-id>" }],
  "workflows": [{ "name": "agent-run", "binding": "AGENT_RUN", "class_name": "AgentRun" }],
}

Workflows change how you write the loop. The rules of Workflows page has the full list. These four matter most for an agent:

  • Code outside step.do runs again on resume. When an instance restarts, run() starts from the top and completed steps return their saved results instead of running again. That is why messages can live outside the steps: it gets rebuilt from saved results. Anything outside a step has to be deterministic, so no Date.now() and no random IDs there.
  • Step names must be stable. model 3 has to mean the same step on every replay. Names built from the loop counter are fine; names built from the current time are not.
  • Steps can run more than once. A step that fails partway, or times out, gets retried. Make side effects idempotent. A plain UPDATE ... WHERE id = ? is safe to repeat. An INSERT without a unique key isn't.
  • Step results are capped at 1 MB. If a tool returns a large document, write it to R2 or D1 inside the step and return a key.

Each step also brings its own retry policy and timeout, which is most of what the deadline and budget logic from my earlier post had to build by hand. You still need the turn cap and a spend limit.

Why not just waitUntil?

ctx.waitUntil() keeps a Worker alive after it responds, but the docs give it at most 30 seconds after the response. An agent run can easily take longer, and if it gets cut off there's no retry and no record of how far it got. Use waitUntil for small fire-and-forget jobs like logging. Use a Workflow for the agent.

D1 limits that matter for agents

  • One database handles one query at a time. The docs say each D1 database is single-threaded. Short indexed queries can do roughly a thousand per second. Slow queries cut that a lot. Keep queries small and indexed, and don't store transcripts in a column you scan.
  • Each database can hold up to 10 GB on the paid plan. The docs suggest splitting data across many databases, for example one per tenant, instead of growing one big one. Up to 50,000 databases per account is allowed on Workers Paid.
  • A query takes at most 100 bound parameters, and a row can be at most 2 MB. Batch inserts need chunking, and large blobs belong in R2.
  • It's SQLite. There's no LISTEN/NOTIFY, no Postgres extensions, no row-level security. Your app does the scoping.

When I wouldn't use this

  • Heavy write traffic on shared data. Many users writing to the same rows at high rates will run into D1's single-threaded design. Postgres handles that better.
  • You need Postgres features. pgvector with a large index, complex joins over big tables, or existing tooling. You can reach an existing Postgres from Workers through Hyperdrive, which is often the better answer than migrating to D1.
  • CPU-heavy work. Parsing huge PDFs or running local models won't fit in 5 minutes of CPU per invocation. Run that in a container or on a VM and call it as a tool.
  • Node-only dependencies. The nodejs_compat flag covers a lot of Node, but not everything. Check your SDKs early.
  • Live, stateful sessions. If the user is chatting in real time over a WebSocket, use a Durable Object (or the Agents SDK) for the session, and D1 for the records you query later.

Takeaways

  • Accept the request in a Worker, store state in D1, and run the agent loop in a Workflow.
  • Make each model call and each tool call its own step. You get saved progress, retries and timeouts for free.
  • Keep code outside steps deterministic, keep step names stable, and make side effects safe to repeat.
  • Design D1 tables around short, indexed queries. Put large blobs in R2 and shard by tenant when you outgrow one database.
  • Choose Postgres, containers or Durable Objects when the workload calls for them. Mixing is fine.
© 2026 Awsaf Alam