What is AI Engineering
Most solo operators hear “AI engineering” and picture a team of PhDs fine‑tuning foundation models. In reality, it’s the practice of stitching together prompts and APIs, plus tiny automation blocks, to get reliable output without a ML PhD. After reading this, you’ll be able to sketch a working AI‑driven workflow, spot where it will break, and decide whether to build it yourself or grab a pre‑made blueprint.
What most guides get wrong about AI engineering
Many tutorials treat AI engineering as a prompt‑crafting contest. They tell you to write longer, cleverer prompts and assume the model will obey. I think that misses the point: the model is a stochastic component, not a deterministic function. The real work lies in validation, fallback logic, and observability.
One common mistake is to skip input sanitization. If you feed raw user text into a language model, you risk prompt injection or unexpected token costs. A simple length check and a deny‑list of disallowed phrases can save you from surprise bills.
Another oversight is treating the model’s output as final. In production, you need a verification step—maybe a regex, a schema check, or a second model call—to catch hallucinations before they reach your customer.
It’s simpler than you think.
How do you handle hallucinations when the model drifts?
Hallucinations appear when the model’s internal distribution shifts, often after a provider updates its backend. The first sign is a sudden rise in nonsensical answers or fabricated facts. You need a lightweight detector that runs on every response.
One approach is to ask the model to self‑grade its confidence. Add a line to your prompt like: “On a scale of 0‑100, how confident are you that the above answer is factually correct? Return only the number.” If the score falls below a threshold, you trigger a fallback—perhaps a rule‑based answer or a notification to review.
Another tactic is to embed a canonical fact set. For a cold‑email pipeline, you might keep a list of known company names and domains. After the model generates a line, you check whether any token matches a known bad pattern (e.g., a made‑up TLD like .xyz that you never allow). If it does, you discard the output and retry with a lower temperature.
These checks add a few milliseconds but save you from sending gibberish to prospects.
A concrete named example: building a cold‑email pipeline with GPT‑4 and Make.com
Let’s walk through a real workflow I run for a freelance outreach agency. The goal: take a list of LinkedIn URLs, pull the person’s name and company, generate a personalized first line, and send it via Gmail.
First, we use Make.com to scrape LinkedIn public pages via a simple HTTP GET (no login needed because we only scrape public profiles). The scenario starts with a Webhook that receives a CSV upload.
Next, we pass the scraped HTML to a Text Parser module that extracts the name and company using CSS selectors. This step is deterministic and cheap.
Then we call the OpenAI API with GPT‑4. Here’s the prompt I use, stored as a constant in the scenario:
You are a friendly sales assistant. Given the prospect's name {{name}} and company {{company}}, write a single sentence that compliments a recent achievement of the company and suggests a brief call. Keep it under 20 words. Do not make up facts.
The temperature is set to 0.3 to reduce creativity. The max tokens is 60.
After the API returns the line, we run a quick validation: we check that the output contains the prospect’s name and does not include any of the blocked phrases [“I am an AI”, “As an AI language model”]. If it fails, we retry once with a lower temperature (0.1).
Finally, we use the Gmail module to send the email. The whole scenario runs in under two seconds per lead.
I love how Make.com’s scenario execution history lets you replay failures with a click—no need to dig through logs.
