Tutorial

How to Build an AI Workflow: A Step-by-Step Guide for Your First Automation

Build your first AI workflow step by step: pick the job, draw the steps, choose the trigger, place the AI step, add a human gate, test, and run it on a schedule.

Contents

A straight line of white dominoes receding into a dark background

Image: Tom Wilson on Unsplash.

An AI workflow is a sequence of steps you define in advance, where one or more of the steps uses a language model to read, judge or write. The rest of the steps are ordinary automation: a trigger, some data fetches, a branch, a write.

That split is the whole trick. Anthropic defines workflows as "systems where LLMs and tools are orchestrated through predefined code paths," and recommends them for tasks that can be "easily and cleanly decomposed into fixed subtasks" (Anthropic, 2024). You keep the predictability of a script and put the model only where rules run out.

This tutorial builds one real workflow from scratch: an inbox triage that labels incoming support email, drafts replies to routine questions and escalates the rest. The steps apply to any tool, and the last section shows how the same build looks in AGNT. If you want an open-ended agent rather than a fixed path, read How to Build an AI Agent instead.

Step 1: Pick a job that repeats and has a clear "done"

Good first workflows share three traits:

  • They happen often. Daily or weekly, so the time saved compounds.
  • They have a checkable output. A label, a drafted reply, a filled row.
  • Mistakes are cheap or catchable. A wrong draft that a person reviews costs little; a wrong refund does not.

OpenAI's agent guide points at the same targets: work with "heavy reliance on unstructured data," such as reading documents or interpreting natural language, where rule-based automation has struggled (OpenAI guide). Email triage fits: the input is messy text, the output is a category and a draft.

Stuck for ideas? We keep a list of 100 AI automation ideas with triggers, outputs and risk levels.

Step 2: Write the steps down before you open any tool

Describe the job as a numbered list a new colleague could follow. For our triage:

  1. When a new email arrives in the support inbox,
  2. read the subject and body,
  3. decide which category it belongs to: billing, bug, how-to or other,
  4. if it is a how-to question, draft a reply from the help docs,
  5. if it is billing or a bug, flag it for a person with a one-line summary,
  6. log what happened.

Now mark each step as rule or judgment:

Step Rule or judgment? Why
1. New email arrives Rule A trigger, no thinking required
2. Read email Rule A data fetch
3. Categorise Judgment Natural language, fuzzy boundaries
4. Draft reply Judgment Writing grounded in docs
5. Flag for a person Rule Once the category is known, routing is fixed
6. Log Rule A write

Two judgment steps out of six is typical. Anthropic calls this shape routing: classify an input, then send it to a specialised follow-up. It "works well for complex tasks where there are distinct categories that are better handled separately" (Anthropic, 2024).

Step 3: Choose the trigger

Every workflow starts from something that happens. The common trigger types are:

  • Schedule: every hour, every Monday at 9:00.
  • Webhook: another system calls your workflow's URL when an event happens, such as a form submission or a payment.
  • Email or message received: a new message in an inbox or channel.
  • Manual: you press run. Use this while testing.

Start with manual, even if the real trigger will be an email. You want to feed the same test email in ten times while you tune the AI step.

Step 4: Build the rule steps first

Wire up the plumbing without any AI:

  1. Trigger → fetch the email → write a placeholder category of "other" → log.
  2. Run it. Confirm the email text arrives intact and the log line is written.

Plumbing bugs, such as wrong credentials, empty fields or encoding problems, are much easier to find before a model is involved. Once the skeleton runs end to end, every later failure is about the AI step.

Step 5: Add the AI step with a tight contract

Replace the placeholder with a model call. The prompt should specify three things: the input, the allowed outputs, and what to do when unsure.

You categorise customer support emails.

Input: the email subject and body below.
Output: JSON with exactly these fields:
  {"category": "billing" | "bug" | "how_to" | "other",
   "confidence": number between 0 and 1,
   "summary": one sentence, max 25 words}

If the email fits none of the categories, or you are unsure, use "other".
Do not invent order numbers, names or facts that are not in the email.

Subject: {{email.subject}}
Body: {{email.body}}

A fixed set of labels and a structured output make the next step, a plain branch on category, deterministic. That is the point: the model's answer feeds a rule, and the rule decides what happens.

For the drafting step, give the model the relevant help-doc text rather than asking it to recall your product from memory. Retrieval grounds the answer in your docs; see RAG.

Step 6: Put a human gate where a mistake would cost you

Decide in advance which outputs go straight out and which wait for a person. For a first workflow, the answer is usually nothing customer-facing goes out unreviewed. The workflow drafts; a person approves.

OpenAI's guidance names the two situations that should trigger human review: exceeding failure thresholds, and high-risk actions that are "sensitive, irreversible, or have high stakes" (OpenAI guide). You can loosen the gate later, category by category, once the logs show the drafts are reliably good. See human-in-the-loop.

Step 7: Test with a small, real sample

Pull 15 to 20 real past emails, including awkward ones: angry, multilingual, two questions in one, spam. Run each through the workflow and check the category by hand.

You do not need a big test set to start. Anthropic's research team began evaluating with about 20 queries, because early on "a prompt tweak might boost success rates from 30% to 80%," and changes that large show up in a small sample (Anthropic, 2025).

Keep the sample. Every time you change the prompt, re-run it and check nothing that used to pass now fails.

Step 8: Turn on the real trigger and read the logs

Switch from manual to the email trigger and let it run for a week with the human gate on. Each day, skim the log:

  • Which categories get overridden by the reviewer most often?
  • Are there emails consistently marked "other" that deserve their own category?
  • How long does each run take, and what does it cost?

This is the observe-and-improve loop. It is also why the per-step log matters more than the prompt: without it, you cannot tell a bad category from a bad fetch.

Step 9: Version it before you change it

Once the workflow is working, save a version before each edit. When a "small prompt tweak" breaks something, you want to roll back in one step rather than reconstruct yesterday's prompt from memory.

Common mistakes

  • Letting the model do rule work. If a step can be an if, make it an if. It is faster, cheaper and never hallucinates.
  • Free-text outputs feeding branches. "Probably billing?" breaks a branch. Use a fixed label set and structured output.
  • No "unsure" path. Without an explicit fallback, the model will guess confidently.
  • Testing only happy paths. Your real inbox contains the weird email on day one.
  • Removing the human gate too early. Remove it per category, only after the logs justify it.

Building it in AGNT

Disclosure: AGNT is our product. In AGNT the same build is a workflow on a canvas:

  1. Drag an Email trigger (or Manual while testing) onto the canvas.
  2. Add an Agent or AI generation node with the categorisation prompt above.
  3. Add a Branch node on category.
  4. On the how-to branch, add a node that drafts the reply; on the billing and bug branches, a node that sends the summary to Slack or your tracker.
  5. Run it. Each node records its input, output and duration, so a bad result points to the exact step.

Workflows are versioned automatically, with named checkpoints and one-step rollback. You can also describe the workflow in chat and have Annie build it, then edit the result on the canvas. It all runs on your machine, with no per-task meter; Community Core is free. The Getting Started guide takes about ten minutes, and a ready-made email triage workflow is in the marketplace if you would rather fork than build.

FAQ

What is the difference between an AI workflow and an AI agent?

In a workflow you define the steps and the model fills in judgment at specific points. In an agent the model decides the steps itself. Workflows are more predictable; agents handle tasks whose steps you cannot know in advance. See AI agents vs workflows.

Do I need to code to build an AI workflow?

No. Visual builders let you connect triggers, AI steps and actions on a canvas. Knowing basic logic, such as branches, loops and structured data, helps more than knowing a programming language.

How much does an AI workflow cost to run?

Two costs: the platform (per task, per execution, or flat) and the model tokens. A short classification call costs little per run. Measure a week of real runs before you estimate annual cost.

What is the best first AI workflow to build?

Something frequent, text-heavy and reviewable: inbox triage, meeting notes to action items, or a weekly report assembled from tools you already use.

Sources

AI WorkflowWorkflow AutomationAI AutomationTutorialGetting Started