Duck
AI Tools7 min read

ai system design

Samet Turan— Editor··7 min read

Learn how to design reliable AI systems for solo operators, from prompt chaining to error handling, with real examples and a ready-to-use blueprint.

ai system design

Most solo operators hit a wall when they try to turn a clever ChatGPT prompt into a repeatable workflow. The prompt works great in the chat box, but once you wrap it in a script or a no‑code tool, edge cases pop up and the whole thing falls apart. After reading this, you’ll be able to sketch a modular AI system, add simple guards, and deploy a working version using the blueprint we’ve packaged at /vault.

What most guides get wrong about AI system design

Many tutorials start with a shiny diagram of agents talking to each other and then jump straight to code. They skip the boring part: defining clear contracts between each step. If you don’t know what output format the first prompt promises, the second step will choke on unexpected JSON or plain text. I’ve seen guides that tell you to “just add more context” and call it a day, which leaves you debugging hallucinations at 2 a.m.

The real mistake is treating the AI like a black box that magically understands intent. In practice you need to treat each call as a function with typed inputs and outputs. That means you spend time writing a short schema—maybe just a JSON‑Schema snippet—and you validate it before passing data downstream. It feels tedious, but it saves hours of rework later.

Another common slip is over‑engineering the orchestration layer. You don’t need a full‑blown workflow engine with queues and retries for a solo‑person project. A simple linear script that checks each step and logs failures is enough. The guides that push you toward complex state machines are solving a problem you don’t have yet.

How do you decide where to put humans in the loop?

You put a human in the loop when the cost of a mistake is higher than the time it takes to review. For a cold‑email pipeline, a wrong name or company isn’t catastrophic; you can let the AI run and then skim the output. For a legal‑draft generator, a single incorrect clause could create liability, so you want a human to approve before sending.

I like to draw a quick risk matrix: impact (low, medium, high) on the vertical axis and frequency (rare, occasional, frequent) on the horizontal. Anything in the high‑impact cell gets a manual check, even if it happens rarely. Low‑impact, frequent items stay fully automated.

One concrete example: I run a lead‑enrichment scraper that pulls company size from LinkedIn. The impact of a wrong size is low—it just means a slightly off‑target email—so I let it run unattended. But when the same scraper pulls a contact’s direct phone number, I flag any result that doesn’t match a known pattern and send it to me for a quick glance. That keeps the false‑positive rate under 5 % without slowing the pipeline.

(which, yes, is annoying) but it’s the only way I’ve found to keep trust in the system.

Building a reliable prompt chain: a concrete example

Let’s walk through a simple three‑step chain that turns a product idea into a one‑page landing‑page copy. We’ll use the OpenAI API, but the same logic works with Claude or any other LLM.

  1. Step 1 – Idea expansion: Prompt: “You are a product strategist. Given the core idea ‘{idea}’, list three customer pain points and three possible features that address each pain. Return JSON with keys ‘pain_points’ and ‘features’.”
  2. Step 2 – Feature selection: Prompt: “You are a UX designer. From the features list, pick the two that deliver the highest value with the lowest effort. Return JSON with key ‘selected_features’.”
  3. Step 3 – Copy generation: Prompt: “You are a copywriter. Write a headline, a sub‑headline, and three bullet points for a landing page that highlights the selected features. Use a friendly, benefit‑first tone. Return JSON with keys ‘headline’, ‘subheadline’, ‘bullets’ (array of three strings).”

Each step validates its output with a tiny JSON‑Schema check. If validation fails, the script logs the error, waits two seconds, and retries the same step up to three times. After three failures it sends me a Slack alert and stops.

Here’s what the validation looks like in Python‑like pseudocode (you can copy‑paste into a Make (formerly Integromat) or n8n HTTP module):

if not jsonschema.validate(output, schema):
log.error(f"Validation failed at step {step}")
if retry < MAX_RETRIES: time.sleep(2) retry += 1 continue else: alert.send(f"Step {step} stuck after {MAX_RETRIES} tries") break

The cost for this chain is tiny: each call uses about 800 tokens, so at $0.002 per 1k tokens you’re looking at roughly $0.005 per run. Even if you run it 200 times a day, that’s under $3 a month.

How to debug when your AI system breaks

When something goes wrong, the first place to look is the log file. I keep a simple JSON line for every call: timestamp, step name, prompt hash, raw response, validation status, and any error message. Grepping for "validation failed" instantly shows you where the chain stalled.

If the logs look fine but the output is nonsense, check the prompt hash. A tiny change in whitespace or a stray character can change the model’s behavior. I store the exact prompt string that was sent, so I can diff it against the version in my repo.

Sometimes the model returns a valid JSON but with the wrong data type—like a string where you expected a number. That’s why I always coerce types after validation: int(field) or float(field) wrapped in a try‑except that logs a warning and substitutes a default.

One gripe I have with the OpenAI API is the occasional silent rate‑limit return: you get a 200 OK with an empty body and no error message. It’s frustrating because your retry loop thinks everything is fine and moves on. I added a check that if the response length is under 50 bytes, I treat it as a failure and trigger a retry.

Finally, if the system works locally but fails in the cloud, look at environment variables. I once spent an hour debugging why the API key was empty—only to discover that the deployment platform masks secrets in the UI but still injects them at runtime. A simple `print(os.getenv('OPENAI_KEY'))` (removed in production) revealed the issue.

Cost and tooling: what you actually need to pay for

You don’t need a fancy orchestration platform to get started. A cheap VPS ($5/mo) or even a free tier of Render or Fly.io will run a Python script that hits the LLM APIs. The real cost comes from the API calls themselves.

For the example chain above, each run is about $0.005. If you run it 1000 times a month, that’s $5. Add a modest logging service like Logtail ($9/mo) and you’re still under $20. I’ve seen guides push you toward a $199/mo "enterprise" agent framework that includes a dashboard you’ll never use. Honestly, that’s ridiculous for what you get.

On the flip side, I love the built‑in token counter in the OpenAI library. It lets me predict the bill before I even hit “run”. I set a daily limit of $10 in my script; if the projected cost exceeds that, the script aborts and sends me a Slack note. That feature alone has saved me from surprise bills more than once.

If you prefer a no‑code route, Make.com’s HTTP module works fine. The free plan gives you 1000 operations a month, which is enough for light testing. The paid plan at $29/mo unlocks unlimited operations and adds error‑handling paths—$29/mo is fair for the reliability you get.

Putting it all together: a simple architecture you can copy

Here’s the layout I use for almost every solo‑operator AI system:

If you want the deep cut on this, deeper coverage of AI agent platforms.

  • Input trigger (webhook, schedule, or manual button)
  • Validation layer – checks that required fields are present
  • Prompt executor – sends the LLM call, logs prompt and response, validates output
  • Error handler – on validation failure, retries up to N times, then alerts
  • Output sink – writes to a Google Sheet, sends an email, or posts to a Slack channel
  • Monitoring – a simple cron job that reads the log file and sends a daily summary of success/failure rates

All of these pieces can be built with a single Python file and a few environment variables. If you’d rather skip the build and deploy a working version in an afternoon, we’ve packaged this workflow as a blueprint at deepusecase.com/vault.

— The Colophon

One AI tool. Tested. Reviewed.
In your inbox every Sunday.

~3 minute read. Real outcomes from operators, not marketers.

Free. One email per Sunday. Unsubscribe in one click.