Skip to main content

Command Palette

Search for a command to run...

From Timeouts To Workflows

Updated
•20 min read•View as Markdown
From Timeouts To Workflows
V
Hey everyone, my name is Ved and I am a passionate and curios developer, Currently learning and sharing my learnings in the best and easiest way possible

You clicked "Review my pull request". The page spun for 90 seconds, then showed a 504 error. ⏳

You build an AI code reviewer. A developer opens a pull request, your server fetches the diff, sends it to an LLM, waits for the review, and posts it back. On small PRs it works beautifully. On a big PR, the request hangs, the browser gives up, and the user sees a timeout, even though your server is still quietly working in the background.

Some work just doesn't fit inside a single web request. Today we'll trace why, starting from that problem, all the way up to background workflows with Inngest, webhooks that kick off AI agents, and what happens when things fail.


The Problem with Doing Everything Inside a Web Request

How a normal request works

In a typical web app, the browser sends a request, your server does some work, and sends a response. The browser waits the whole time.

Browser ──▶ Server (does the work...) ──▶ Response
              ▲
              └── the user is staring at a spinner meanwhile

For quick things like fetching a profile or saving a form, this is perfect.

Requests have limits

Somewhere between the browser and your code, there are timers running.

  • Browsers and clients give up after a while

  • Load balancers and proxies cut connections that stay open too long

  • Serverless platforms often cap how long a single function can run

  • Webhook senders like GitHub expect a quick reply (GitHub waits around 10 seconds)

The exact numbers differ by platform, but the lesson is the same: a request is meant to be short.

What goes wrong with long work

  • Timeouts: the work gets cut off halfway

  • Lost work: if the server restarts mid-request, everything in progress vanishes

  • A frozen user: nobody wants to watch a spinner for two minutes

  • No retry: if step 4 of 6 fails, you start over from the beginning, or not at all

  • Wasted money: imagine paying for an LLM call, then losing the result to a timeout

A simple test

Ask yourself: "Could this take more than a few seconds, or could it fail halfway?" If yes, it probably shouldn't live inside a web request.


Long-Running AI Tasks

AI makes this problem much worse

Traditional apps mostly do quick database reads. AI apps do things that are slow and unpredictable by nature.

  • An LLM call alone can take from a few seconds to a minute or more

  • An agent might make ten or twenty LLM and tool calls in a row

  • Each call depends on the previous one, so they can't all run at once

  • The total time isn't even known in advance, because the agent decides how many steps it needs

Examples of tasks that take too long

  • Code review: fetch a large diff, split it into files, review each one, and combine the feedback

  • Document processing: extract text from a 200-page PDF, chunk it, embed it, summarize it

  • Agent execution: research a topic by searching, reading pages, and writing a report

  • Data enrichment: process thousands of rows, calling an LLM on each one

  • Media work: transcribing or generating audio, images, or video

The timing, in numbers

Fetch the PR diff            ██                      ~2s
Summarize each file (x12)    ████████████████        ~40s
Check for security issues    ██████████              ~20s
Write the final review       █████                   ~10s
Post the comment             █                       ~1s
                             ─────────────────────────────
                             Total ≈ 73 seconds   😬

That's far beyond what any one request should hold open. We need a different shape for this kind of work.


What Background Workflows Are

The idea

Instead of doing the work while the user waits, we do this:

  1. Accept the request and say, "Got it, working on it!" immediately

  2. Do the actual work in the background, separately

  3. Deliver the result when it's ready

The restaurant analogy

At a restaurant, the waiter doesn't stand at your table for 20 minutes while the chef cooks. They take your order, hand you a token, and go. The kitchen works on it in the background. When it's ready, they bring it over.

Your request is the order. The kitchen is the background worker.

Synchronous request vs background workflow

SYNCHRONOUS (everything inside the request)

User ──▶ Server ──▶ [ step 1 → step 2 → step 3 → step 4 ] ──▶ Response
              user waits the entire time, timeout risk 😬


BACKGROUND WORKFLOW

User ──▶ Server ──▶ "Accepted! ✅" (instant response)
              │
              └──▶ Queue ──▶ Worker runs [ step 1 → 2 → 3 → 4 ]
                                      │
                                      ▼
                         Result saved / user notified

Side by side

Synchronous Request Background Workflow
User waits for The whole task Just an acknowledgement
Timeout risk High Low
If it fails halfway Usually start over Can retry just the failed step
Good for Quick, small tasks Long, multi-step, or unreliable tasks
Result delivery In the response Notification, webhook, polling, or live updates
Complexity Simple More moving parts

A workflow is more than a job

A basic background job is "run this function later". A workflow is a series of steps, where each step can be saved, retried, and resumed independently. That distinction matters a lot, as we'll see soon.


Why AI Applications Need Asynchronous Processing

1. LLM calls are slow and flaky

APIs get rate-limited. Providers have brief outages. A call that works nine times might return a 429 or 500 on the tenth. When your task has twenty such calls, something will fail eventually.

2. Every step costs real money

If step 7 of 8 fails and you restart from scratch, you pay for steps 1 to 7 again. For AI work, that waste adds up fast. Saving progress between steps isn't just nice, it's cheaper.

3. Better user experience

The user shouldn't be stuck. They should see "Review in progress..." and get notified when it's done. Async lets the app stay responsive.

4. Scaling and control

Sudden traffic spikes could overwhelm your LLM provider's rate limits. With background processing, you can control how many tasks run at once and spread the load out.

5. Work needs to survive restarts

Deployments happen. Servers restart. A background workflow that saves its progress can pick up where it stopped, instead of losing everything.

The pattern

  • Fast work goes in the request

  • Slow, multi-step, or failure-prone work goes in a workflow

That's what we need a tool for.


What Inngest Does

The problem it solves

You could build all of this yourself. A queue, workers, retry logic, state storage, scheduling, monitoring. It's a lot of plumbing that has nothing to do with your actual product.

Inngest is a platform for durable background workflows. You write your workflow as normal code, and Inngest handles running it reliably: queuing, retries, saved state, concurrency limits, scheduling, and a dashboard to see what happened.

The key idea: durable execution

"Durable" means the workflow survives failures. Each step's result is saved. If something breaks, the workflow continues from the last successful step, not from the beginning.

How you use it

Your code stays in your own app, on your own hosting. You expose a small endpoint, and Inngest calls it whenever a workflow needs to run. Think of it like this:

┌───────────────────┐   triggers a run    ┌──────────────────────┐
│     Inngest       │ ──────────────────▶ │  Your app            │
│ (queue, state,    │                     │  (your functions     │
│  retries, logs)   │ ◀────────────────── │   at /api/inngest)   │
└───────────────────┘   step results      └──────────────────────┘

Inngest decides when to run things and remembers the results. Your code decides what the steps do.

Setting it up

npm install inngest
// src/inngest/client.ts
import { Inngest } from "inngest";

export const inngest = new Inngest({ id: "my-ai-app" });

That's the whole client. It groups the functions in your service under one app ID.

What you get

  • Durable steps with saved results

  • Automatic retries with backoff

  • Event triggers, cron schedules, and delays

  • Concurrency and rate limits to protect your APIs

  • A dashboard showing every run, every step, and every error

  • Works with any framework (Next.js, Express, and others), in TypeScript, Python, and Go


Events and Workflow Execution

Everything starts with an event

In Inngest, workflows are triggered by events. An event is just a small message saying, "something happened", with a name and some data.

await inngest.send({
  name: "app/document.uploaded",
  data: { documentId: "doc_42", userId: "user_7" },
});

You may remember the event-driven thinking from the Kafka article. Same idea: the sender announces what happened, and doesn't need to know who reacts to it.

Functions listen for events

You write a function and tell it which event should trigger it:

// src/inngest/functions.ts
import { inngest } from "./client";

export const processDocument = inngest.createFunction(
  { id: "process-document", triggers: { event: "app/document.uploaded" } },
  async ({ event, step }) => {
    const text = await step.run("extract-text", async () => {
      return await extractText(event.data.documentId);
    });

    const summary = await step.run("summarize", async () => {
      return await summarizeWithLLM(text);
    });

    await step.run("save-summary", async () => {
      await saveSummary(event.data.documentId, summary);
    });

    return { documentId: event.data.documentId, status: "done" };
  }
);

Serving your functions

Inngest needs a way to reach your functions. In a Next.js app, that's one route:

// app/api/inngest/route.ts
import { serve } from "inngest/next";
import { inngest } from "@/inngest/client";
import { processDocument } from "@/inngest/functions";

export const { GET, POST, PUT } = serve({
  client: inngest,
  functions: [processDocument],
});

What is a step?

A step (step.run) is a unit of work with a name. Inngest treats each step as its own saved, retriable chunk:

  • When a step succeeds, its result is saved and it will not run again

  • When a step throws an error, only that step is retried

  • The result of each step is stored as JSON, so you can pass it into later steps

The full journey of an event

Event ──▶ Inngest ──▶ Your function
"document          │
 uploaded"         ├─ step: extract-text   ✅ saved
                   ├─ step: summarize      ✅ saved
                   └─ step: save-summary   ✅ saved
                                │
                                ▼
                         Run complete

Other ways to pause and wait

Workflows aren't limited to running straight through. Inngest also lets a workflow sleep (step.sleep), wait for another event (step.waitForEvent), or call another function. A workflow can pause for hours or days without holding a server open, which is hard to do with plain requests.

Testing locally

Inngest has a local Dev Server where you can send test events and watch each run, its steps, and their outputs.

npx inngest-cli@latest dev

Retries and Reliable Execution

Failure is normal

With AI work, failures are routine: rate limits, network blips, a provider outage, a malformed response. A reliable system doesn't pretend these won't happen. It plans for them.

Step-level retries

When a step throws, Inngest retries that step automatically, with a delay between attempts. The steps before it are not rerun, because their results are already saved.

Attempt 1                              Retry (only the failed step)

step 1: fetch diff       ✅ saved      step 1: fetch diff       ⏭️ skipped (saved)
step 2: LLM review       ❌ 429 error  step 2: LLM review       ✅ success
                                       step 3: post comment     ✅ success

Why this saves money and time

Imagine step 1 was an expensive call and step 2 failed. Without durable steps, you'd pay for step 1 again. With them, only step 2 is retried. The LLM call you already paid for isn't wasted.

How the retry timeline looks

Time ─────────────────────────────────────────────▶

step: call-llm
   try 1 ❌ ──wait──▶ try 2 ❌ ──wait longer──▶ try 3 ✅
                                                  │
                                                  ▼
                                         workflow continues

By default, Inngest retries a failing function several times before giving up, and you can configure that number per function:

inngest.createFunction(
  {
    id: "process-document",
    retries: 5,
    triggers: { event: "app/document.uploaded" },
  },
  async ({ event, step }) => { /* ... */ }
);

Not every error deserves a retry

If the failure is permanent, like "this document is corrupted" or "the user doesn't exist", retrying five times just wastes time. For those, you can throw a special error that tells Inngest to stop retrying:

import { NonRetriableError } from "inngest";

await step.run("validate", async () => {
  if (!doc.isReadable) {
    throw new NonRetriableError("Document is corrupted");
  }
});

What if it still fails?

After all retries are used up, the run is marked as failed, and you can see exactly which step failed, with which error, in the dashboard. You can also add failure handling, like alerting your team or marking the job as failed for the user.

One rule to remember: make steps safe to repeat

Retries mean a step might run more than once. Imagine a step that charges a card or sends an email. If it succeeds on the provider's side but the response gets lost, a retry could do it twice. For steps with side effects, use idempotency, such as an idempotency key or a check like "has this already been done?".


Webhooks Triggering AI Workflows

What a webhook is

A webhook is a message that another service sends to you when something happens. GitHub says, "a pull request was opened". Stripe says, "a payment succeeded". Instead of you constantly checking, they call a URL you provide.

The challenge

Webhook senders expect a fast reply. They usually wait only a few seconds, and may retry or mark the delivery as failed if you're slow. But handling the event might take over a minute of AI work.

The fix is simple: receive fast, process later.

The pattern

  1. The webhook endpoint verifies the message is genuine

  2. It hands the work to a background workflow by sending an Inngest event

  3. It responds immediately with a 200

Example: GitHub PR → AI code review

GitHub PR opened
      │
      ▼
Webhook ──▶ Your endpoint ──▶ inngest.send(event) ──▶ returns 200 instantly
                                      │
                                      ▼
                                   Inngest
                                      │
                                      ▼
                               AI Review Workflow
                                      │
        ┌──────────────┬──────────────┴───────────┬─────────────────┐
        ▼              ▼                          ▼                 ▼
   fetch diff     review files             summarize issues     post comment
                                                                    │
                                                                    ▼
                                                      Review appears on the PR ✅

Step 1: The webhook endpoint

// app/api/github/webhook/route.ts
import crypto from "crypto";
import { inngest } from "@/inngest/client";

export async function POST(req: Request) {
  const body = await req.text();
  const signature = req.headers.get("x-hub-signature-256") ?? "";

  // Verify the request really came from GitHub
  const expected =
    "sha256=" +
    crypto
      .createHmac("sha256", process.env.GITHUB_WEBHOOK_SECRET!)
      .update(body)
      .digest("hex");

  const valid =
    signature.length === expected.length &&
    crypto.timingSafeEqual(Buffer.from(signature), Buffer.from(expected));

  if (!valid) {
    return new Response("Invalid signature", { status: 401 });
  }

  const payload = JSON.parse(body);

  // Only care about newly opened pull requests
  if (req.headers.get("x-github-event") === "pull_request" &&
      payload.action === "opened") {
    await inngest.send({
      name: "github/pull_request.opened",
      data: {
        repo: payload.repository.full_name,
        prNumber: payload.pull_request.number,
      },
    });
  }

  // Respond immediately, the heavy work happens in the background
  return new Response("ok", { status: 200 });
}

Notice the signature check. Never skip it. Without it, anyone who finds your URL could trigger your AI workflow, and your LLM bill.

Step 2: The review workflow

import Anthropic from "@anthropic-ai/sdk";
import { inngest } from "./client";

const anthropic = new Anthropic(); // reads ANTHROPIC_API_KEY

export const reviewPullRequest = inngest.createFunction(
  { id: "review-pull-request", triggers: { event: "github/pull_request.opened" } },
  async ({ event, step }) => {
    const { repo, prNumber } = event.data;

    // 1. Get the diff
    const diff = await step.run("fetch-diff", async () => {
      return await getPullRequestDiff(repo, prNumber);
    });

    // 2. Ask the LLM to review it
    const review = await step.run("llm-review", async () => {
      const response = await anthropic.messages.create({
        model: "claude-sonnet-5-5",
        max_tokens: 2000,
        messages: [
          {
            role: "user",
            content: `Review this pull request diff. Point out bugs, risks, and improvements:\n\n${diff}`,
          },
        ],
      });
      const block = response.content[0];
      return block.type === "text" ? block.text : "";
    });

    // 3. Post the result back to the PR
    await step.run("post-comment", async () => {
      await postPullRequestComment(repo, prNumber, review);
    });

    return { repo, prNumber, status: "reviewed" };
  }
);

Here getPullRequestDiff and postPullRequestComment are your own helper functions, perhaps using GitHub's Octokit library. If the LLM call hits a rate limit, only llm-review retries. The diff is already saved.


Agents Inside Background Workflows

Agents are the perfect fit

Remember the agent loop from earlier in this series: perceive, decide, act, observe, repeat. An agent might loop ten or twenty times, calling tools along the way. That's exactly the kind of long, unpredictable, failure-prone work that belongs in a background workflow.

Make each action a step

The trick is to wrap each model call and each tool call in its own step.run. Every iteration of the loop becomes a checkpoint.

export const researchAgent = inngest.createFunction(
  { id: "research-agent", triggers: { event: "app/research.requested" } },
  async ({ event, step }) => {
    const MAX_STEPS = 8;                       // a guardrail, remember?
    const messages = [{ role: "user", content: event.data.question }];

    for (let i = 0; i < MAX_STEPS; i++) {
      // DECIDE: ask the model what to do next
      const decision = await step.run(`decide-${i}`, async () => {
        return await callModel(messages);      // returns a tool request or a final answer
      });

      if (decision.type === "final_answer") {
        await step.run("save-result", () => saveResult(event.data.id, decision.text));
        return { status: "done" };
      }

      // ACT: run the requested tool
      const toolResult = await step.run(`tool-${i}`, async () => {
        return await runTool(decision.tool, decision.args);
      });

      // OBSERVE: add the result and loop again
      messages.push({ role: "assistant", content: JSON.stringify(decision) });
      messages.push({ role: "user", content: JSON.stringify(toolResult) });
    }

    return { status: "stopped_at_step_limit" };
  }
);

Here callModel and runTool are placeholders for your own LLM and tool code, like the tool calling we covered earlier.

Why this is powerful

  • Crash on iteration 6? The first 5 iterations are saved, so the agent resumes at 6 instead of starting from zero

  • A tool fails? Only that tool call retries

  • Every decision is visible in the dashboard, so debugging an agent becomes much easier

  • The loop limit still applies, so a confused agent can't run forever

A note on frameworks

Inngest also offers AgentKit, a library for building agents and multi-agent networks that run on top of this durable execution. You can also use Inngest alongside whichever agent framework you already like, and just wrap its calls in steps.

Long waits and humans in the loop

Remember the human approval guardrail for risky actions? A workflow can pause and wait for an approval event without keeping anything running.

const approval = await step.waitForEvent("wait-for-approval", {
  event: "app/refund.approved",
  timeout: "24h",
  match: "data.refundId",
});

if (!approval) {
  // nobody approved within 24 hours
  return { status: "expired" };
}

The workflow sleeps for up to a day, using no resources, then continues when the approval arrives.


Real-World Examples of AI Workflows

1. AI code review bot

A webhook fires when a PR opens. The workflow fetches the diff, reviews each file, summarizes the findings, and posts a comment. Big PRs just take longer, and nobody's waiting on a spinner.

2. Document and PDF processing

A user uploads a 300-page contract. The workflow extracts the text, splits it into sections, summarizes each in parallel, extracts key terms, and emails a report. If one section's summary fails, only that part retries.

Upload ──▶ extract text ──▶ chunk ──▶ ┌─ summarize chunk 1 ─┐
                                      ├─ summarize chunk 2 ─┼──▶ combine ──▶ notify user
                                      └─ summarize chunk 3 ─┘
                                              (parallel steps)

3. Customer support triage

A new ticket arrives. A workflow classifies it, looks up the customer, drafts a reply, and routes it to the right team. If it's a refund above a certain amount, it waits for human approval before acting.

4. Research and report agents

A user asks for a market research report. An agent searches, reads, takes notes, and writes sections over several minutes, saving progress after each step. The user comes back later to a finished report.

5. Content and media pipelines

Transcribe a video, generate chapters and a summary, create social posts, and schedule them for tomorrow morning. Delays and scheduling are natural in a workflow.

6. Scheduled AI jobs

Every Monday at 9 AM, generate a weekly digest of everything that happened across a team's tools, and send it out. This is a cron-style trigger instead of an event.

7. Bulk data enrichment

Run an LLM over 50,000 CRM records. Concurrency limits keep you under the provider's rate limits, and failed records retry on their own.

Where it all fits in the architecture

┌────────────────────────────────────────────────────────────┐
│                     Web App / API                          │
│        (fast requests: auth, reads, accept & enqueue)      │
└──────────────┬─────────────────────────────▲───────────────┘
               │ events / webhooks           │ results, status
               ▼                             │
┌──────────────────────────────┐   ┌─────────┴──────────────┐
│       Inngest                │   │   Database / Storage   │
│  queue · state · retries     │   │   (results, history)   │
└──────────────┬───────────────┘   └─────────▲──────────────┘
               │ runs steps                  │
               ▼                             │
┌────────────────────────────────────────────┴───────────────┐
│               Background Workflows (your code)             │
│     AI agents · LLM calls · tools · MCP servers · memory   │
└────────────────────────────────────────────────────────────┘

Notice how this ties the whole series together. The agents we built, the tools they call (maybe through MCP), and the memory they use all run inside these workflows. The workflow layer is what makes them reliable in production.

When you don't need a workflow

Keep it honest: if your task is a single quick LLM call that returns in two seconds, a normal request is fine. Workflows add moving parts. Reach for them when work is long, multi-step, failure-prone, or needs to wait.


Quick Recap

  • Web requests are meant to be short. Timeouts, restarts, and impatient users make long work inside a request fragile.

  • AI tasks are slow, multi-step, and flaky, so they hit these limits constantly.

  • A background workflow accepts the request instantly, does the work separately, and delivers the result later. It's a series of steps, each saved and retriable.

  • AI apps need asynchronous processing for reliability, cost control, better user experience, and surviving restarts.

  • Inngest runs durable workflows for you, handling queuing, retries, state, concurrency, and visibility, while your code stays in your own app.

  • Events trigger workflows. Steps (step.run) save their results, so successful work is never repeated.

  • When a step fails, only that step retries. Use non-retriable errors for permanent failures, and make steps safe to repeat.

  • Webhooks should be received fast and handed to a workflow. Always verify the signature.

  • Agents fit naturally inside workflows: each decision and tool call becomes a checkpointed step, with a loop limit as a guardrail.

  • Production AI apps use workflows for code review, document processing, support triage, research agents, scheduled jobs, and bulk enrichment.

A web request answers a question. A workflow finishes a job. When your AI work takes longer than a conversation should, stop making the user wait, and let the work run reliably in the background.