How to Build an AI Workflow: A Step-by-Step Guide for Your First Automation
Build your first AI workflow step by step: pick the job, draw the steps, choose the trigger, place the AI step, add a human gate, test, and run it on a schedule.
Contents
- Step 1: Pick a job that repeats and has a clear "done"
- Step 2: Write the steps down before you open any tool
- Step 3: Choose the trigger
- Step 4: Build the rule steps first
- Step 5: Add the AI step with a tight contract
- Step 6: Put a human gate where a mistake would cost you
- Step 7: Test with a small, real sample
- Step 8: Turn on the real trigger and read the logs
- Step 9: Version it before you change it
- Common mistakes
- Building it in AGNT
- FAQ
- Sources
Image: Tom Wilson on Unsplash.
An AI workflow is a sequence of steps you define in advance, where one or more of the steps uses a language model to read, judge or write. The rest of the steps are ordinary automation: a trigger, some data fetches, a branch, a write.
That split is the whole trick. Anthropic defines workflows as "systems where LLMs and tools are orchestrated through predefined code paths," and recommends them for tasks that can be "easily and cleanly decomposed into fixed subtasks" (Anthropic, 2024). You keep the predictability of a script and put the model only where rules run out.
This tutorial builds one real workflow from scratch: an inbox triage that labels incoming support email, drafts replies to routine questions and escalates the rest. The steps apply to any tool, and the last section shows how the same build looks in AGNT. If you want an open-ended agent rather than a fixed path, read How to Build an AI Agent instead.
Step 1: Pick a job that repeats and has a clear "done"
Good first workflows share three traits:
- They happen often. Daily or weekly, so the time saved compounds.
- They have a checkable output. A label, a drafted reply, a filled row.
- Mistakes are cheap or catchable. A wrong draft that a person reviews costs little; a wrong refund does not.
OpenAI's agent guide points at the same targets: work with "heavy reliance on unstructured data," such as reading documents or interpreting natural language, where rule-based automation has struggled (OpenAI guide). Email triage fits: the input is messy text, the output is a category and a draft.
Stuck for ideas? We keep a list of 100 AI automation ideas with triggers, outputs and risk levels.
Step 2: Write the steps down before you open any tool
Describe the job as a numbered list a new colleague could follow. For our triage:
- When a new email arrives in the support inbox,
- read the subject and body,
- decide which category it belongs to: billing, bug, how-to or other,
- if it is a how-to question, draft a reply from the help docs,
- if it is billing or a bug, flag it for a person with a one-line summary,
- log what happened.
Now mark each step as rule or judgment:
| Step | Rule or judgment? | Why |
|---|---|---|
| 1. New email arrives | Rule | A trigger, no thinking required |
| 2. Read email | Rule | A data fetch |
| 3. Categorise | Judgment | Natural language, fuzzy boundaries |
| 4. Draft reply | Judgment | Writing grounded in docs |
| 5. Flag for a person | Rule | Once the category is known, routing is fixed |
| 6. Log | Rule | A write |
Two judgment steps out of six is typical. Anthropic calls this shape routing: classify an input, then send it to a specialised follow-up. It "works well for complex tasks where there are distinct categories that are better handled separately" (Anthropic, 2024).
Step 3: Choose the trigger
Every workflow starts from something that happens. The common trigger types are:
- Schedule: every hour, every Monday at 9:00.
- Webhook: another system calls your workflow's URL when an event happens, such as a form submission or a payment.
- Email or message received: a new message in an inbox or channel.
- Manual: you press run. Use this while testing.
Start with manual, even if the real trigger will be an email. You want to feed the same test email in ten times while you tune the AI step.
Step 4: Build the rule steps first
Wire up the plumbing without any AI:
- Trigger → fetch the email → write a placeholder category of "other" → log.
- Run it. Confirm the email text arrives intact and the log line is written.
Plumbing bugs, such as wrong credentials, empty fields or encoding problems, are much easier to find before a model is involved. Once the skeleton runs end to end, every later failure is about the AI step.
Step 5: Add the AI step with a tight contract
Replace the placeholder with a model call. The prompt should specify three things: the input, the allowed outputs, and what to do when unsure.
You categorise customer support emails.
Input: the email subject and body below.
Output: JSON with exactly these fields:
{"category": "billing" | "bug" | "how_to" | "other",
"confidence": number between 0 and 1,
"summary": one sentence, max 25 words}
If the email fits none of the categories, or you are unsure, use "other".
Do not invent order numbers, names or facts that are not in the email.
Subject: {{email.subject}}
Body: {{email.body}}A fixed set of labels and a structured output make the next step, a plain branch on category, deterministic. That is the point: the model's answer feeds a rule, and the rule decides what happens.
For the drafting step, give the model the relevant help-doc text rather than asking it to recall your product from memory. Retrieval grounds the answer in your docs; see RAG.
Step 6: Put a human gate where a mistake would cost you
Decide in advance which outputs go straight out and which wait for a person. For a first workflow, the answer is usually nothing customer-facing goes out unreviewed. The workflow drafts; a person approves.
OpenAI's guidance names the two situations that should trigger human review: exceeding failure thresholds, and high-risk actions that are "sensitive, irreversible, or have high stakes" (OpenAI guide). You can loosen the gate later, category by category, once the logs show the drafts are reliably good. See human-in-the-loop.
Step 7: Test with a small, real sample
Pull 15 to 20 real past emails, including awkward ones: angry, multilingual, two questions in one, spam. Run each through the workflow and check the category by hand.
You do not need a big test set to start. Anthropic's research team began evaluating with about 20 queries, because early on "a prompt tweak might boost success rates from 30% to 80%," and changes that large show up in a small sample (Anthropic, 2025).
Keep the sample. Every time you change the prompt, re-run it and check nothing that used to pass now fails.
Step 8: Turn on the real trigger and read the logs
Switch from manual to the email trigger and let it run for a week with the human gate on. Each day, skim the log:
- Which categories get overridden by the reviewer most often?
- Are there emails consistently marked "other" that deserve their own category?
- How long does each run take, and what does it cost?
This is the observe-and-improve loop. It is also why the per-step log matters more than the prompt: without it, you cannot tell a bad category from a bad fetch.
Step 9: Version it before you change it
Once the workflow is working, save a version before each edit. When a "small prompt tweak" breaks something, you want to roll back in one step rather than reconstruct yesterday's prompt from memory.
Common mistakes
- Letting the model do rule work. If a step can be an
if, make it anif. It is faster, cheaper and never hallucinates. - Free-text outputs feeding branches. "Probably billing?" breaks a branch. Use a fixed label set and structured output.
- No "unsure" path. Without an explicit fallback, the model will guess confidently.
- Testing only happy paths. Your real inbox contains the weird email on day one.
- Removing the human gate too early. Remove it per category, only after the logs justify it.
Building it in AGNT
Disclosure: AGNT is our product. In AGNT the same build is a workflow on a canvas:
- Drag an Email trigger (or Manual while testing) onto the canvas.
- Add an Agent or AI generation node with the categorisation prompt above.
- Add a Branch node on
category. - On the how-to branch, add a node that drafts the reply; on the billing and bug branches, a node that sends the summary to Slack or your tracker.
- Run it. Each node records its input, output and duration, so a bad result points to the exact step.
Workflows are versioned automatically, with named checkpoints and one-step rollback. You can also describe the workflow in chat and have Annie build it, then edit the result on the canvas. It all runs on your machine, with no per-task meter; Community Core is free. The Getting Started guide takes about ten minutes, and a ready-made email triage workflow is in the marketplace if you would rather fork than build.
FAQ
What is the difference between an AI workflow and an AI agent?
In a workflow you define the steps and the model fills in judgment at specific points. In an agent the model decides the steps itself. Workflows are more predictable; agents handle tasks whose steps you cannot know in advance. See AI agents vs workflows.
Do I need to code to build an AI workflow?
No. Visual builders let you connect triggers, AI steps and actions on a canvas. Knowing basic logic, such as branches, loops and structured data, helps more than knowing a programming language.
How much does an AI workflow cost to run?
Two costs: the platform (per task, per execution, or flat) and the model tokens. A short classification call costs little per run. Measure a week of real runs before you estimate annual cost.
What is the best first AI workflow to build?
Something frequent, text-heavy and reviewable: inbox triage, meeting notes to action items, or a weekly report assembled from tools you already use.
Sources
- Anthropic, Building effective agents (December 2024)
- Anthropic, How we built our multi-agent research system (June 2025)
- OpenAI, A practical guide to building agents
- AGNT, Workflows and Getting Started