ai system design
Most solo operators hit a wall when they try to turn a clever ChatGPT prompt into a repeatable workflow. The prompt works great in the chat box, but once you wrap it in a script or a no‑code tool, edge cases pop up and the whole thing falls apart. After reading this, you’ll be able to sketch a modular AI system, add simple guards, and deploy a working version using the blueprint we’ve packaged at /vault.
What most guides get wrong about AI system design
Many tutorials start with a shiny diagram of agents talking to each other and then jump straight to code. They skip the boring part: defining clear contracts between each step. If you don’t know what output format the first prompt promises, the second step will choke on unexpected JSON or plain text. I’ve seen guides that tell you to “just add more context” and call it a day, which leaves you debugging hallucinations at 2 a.m.
The real mistake is treating the AI like a black box that magically understands intent. In practice you need to treat each call as a function with typed inputs and outputs. That means you spend time writing a short schema—maybe just a JSON‑Schema snippet—and you validate it before passing data downstream. It feels tedious, but it saves hours of rework later.
Another common slip is over‑engineering the orchestration layer. You don’t need a full‑blown workflow engine with queues and retries for a solo‑person project. A simple linear script that checks each step and logs failures is enough. The guides that push you toward complex state machines are solving a problem you don’t have yet.
How do you decide where to put humans in the loop?
You put a human in the loop when the cost of a mistake is higher than the time it takes to review. For a cold‑email pipeline, a wrong name or company isn’t catastrophic; you can let the AI run and then skim the output. For a legal‑draft generator, a single incorrect clause could create liability, so you want a human to approve before sending.
I like to draw a quick risk matrix: impact (low, medium, high) on the vertical axis and frequency (rare, occasional, frequent) on the horizontal. Anything in the high‑impact cell gets a manual check, even if it happens rarely. Low‑impact, frequent items stay fully automated.
One concrete example: I run a lead‑enrichment scraper that pulls company size from LinkedIn. The impact of a wrong size is low—it just means a slightly off‑target email—so I let it run unattended. But when the same scraper pulls a contact’s direct phone number, I flag any result that doesn’t match a known pattern and send it to me for a quick glance. That keeps the false‑positive rate under 5 % without slowing the pipeline.
(which, yes, is annoying) but it’s the only way I’ve found to keep trust in the system.
Building a reliable prompt chain: a concrete example
Let’s walk through a simple three‑step chain that turns a product idea into a one‑page landing‑page copy. We’ll use the OpenAI API, but the same logic works with Claude or any other LLM.
- Step 1 – Idea expansion: Prompt: “You are a product strategist. Given the core idea ‘{idea}’, list three customer pain points and three possible features that address each pain. Return JSON with keys ‘pain_points’ and ‘features’.”
- Step 2 – Feature selection: Prompt: “You are a UX designer. From the features list, pick the two that deliver the highest value with the lowest effort. Return JSON with key ‘selected_features’.”
- Step 3 – Copy generation: Prompt: “You are a copywriter. Write a headline, a sub‑headline, and three bullet points for a landing page that highlights the selected features. Use a friendly, benefit‑first tone. Return JSON with keys ‘headline’, ‘subheadline’, ‘bullets’ (array of three strings).”
Each step validates its output with a tiny JSON‑Schema check. If validation fails, the script logs the error, waits two seconds, and retries the same step up to three times. After three failures it sends me a Slack alert and stops.
Here’s what the validation looks like in Python‑like pseudocode (you can copy‑paste into a Make (formerly Integromat) or n8n HTTP module):
if not jsonschema.validate(output, schema):
log.error(f"Validation failed at step {step}")
if retry < MAX_RETRIES:
time.sleep(2)
retry += 1
continue
else:
alert.send(f"Step {step} stuck after {MAX_RETRIES} tries")
break
The cost for this chain is tiny: each call uses about 800 tokens, so at $0.002 per 1k tokens you’re looking at roughly $0.005 per run. Even if you run it 200 times a day, that’s under $3 a month.
