Duck
AI Tools5 min read

What is AI Engineering

Samet Turan— Editor··5 min read

Learn what AI engineering really means for solo operators, with real prompts, tool costs, and debugging tips you can apply today and scale reliably.

What is AI Engineering

Most solo operators hear “AI engineering” and picture a team of PhDs fine‑tuning foundation models. In reality, it’s the practice of stitching together prompts and APIs, plus tiny automation blocks, to get reliable output without a ML PhD. After reading this, you’ll be able to sketch a working AI‑driven workflow, spot where it will break, and decide whether to build it yourself or grab a pre‑made blueprint.

What most guides get wrong about AI engineering

Many tutorials treat AI engineering as a prompt‑crafting contest. They tell you to write longer, cleverer prompts and assume the model will obey. I think that misses the point: the model is a stochastic component, not a deterministic function. The real work lies in validation, fallback logic, and observability.

One common mistake is to skip input sanitization. If you feed raw user text into a language model, you risk prompt injection or unexpected token costs. A simple length check and a deny‑list of disallowed phrases can save you from surprise bills.

Another oversight is treating the model’s output as final. In production, you need a verification step—maybe a regex, a schema check, or a second model call—to catch hallucinations before they reach your customer.

It’s simpler than you think.

How do you handle hallucinations when the model drifts?

Hallucinations appear when the model’s internal distribution shifts, often after a provider updates its backend. The first sign is a sudden rise in nonsensical answers or fabricated facts. You need a lightweight detector that runs on every response.

One approach is to ask the model to self‑grade its confidence. Add a line to your prompt like: “On a scale of 0‑100, how confident are you that the above answer is factually correct? Return only the number.” If the score falls below a threshold, you trigger a fallback—perhaps a rule‑based answer or a notification to review.

Another tactic is to embed a canonical fact set. For a cold‑email pipeline, you might keep a list of known company names and domains. After the model generates a line, you check whether any token matches a known bad pattern (e.g., a made‑up TLD like .xyz that you never allow). If it does, you discard the output and retry with a lower temperature.

These checks add a few milliseconds but save you from sending gibberish to prospects.

A concrete named example: building a cold‑email pipeline with GPT‑4 and Make.com

Let’s walk through a real workflow I run for a freelance outreach agency. The goal: take a list of LinkedIn URLs, pull the person’s name and company, generate a personalized first line, and send it via Gmail.

First, we use Make.com to scrape LinkedIn public pages via a simple HTTP GET (no login needed because we only scrape public profiles). The scenario starts with a Webhook that receives a CSV upload.

Next, we pass the scraped HTML to a Text Parser module that extracts the name and company using CSS selectors. This step is deterministic and cheap.

Then we call the OpenAI API with GPT‑4. Here’s the prompt I use, stored as a constant in the scenario:

You are a friendly sales assistant. Given the prospect's name {{name}} and company {{company}}, write a single sentence that compliments a recent achievement of the company and suggests a brief call. Keep it under 20 words. Do not make up facts.

The temperature is set to 0.3 to reduce creativity. The max tokens is 60.

After the API returns the line, we run a quick validation: we check that the output contains the prospect’s name and does not include any of the blocked phrases [“I am an AI”, “As an AI language model”]. If it fails, we retry once with a lower temperature (0.1).

Finally, we use the Gmail module to send the email. The whole scenario runs in under two seconds per lead.

I love how Make.com’s scenario execution history lets you replay failures with a click—no need to dig through logs.

How to debug when this breaks

When the pipeline stops delivering emails, the first place to look is the scenario’s execution log in Make.com. Each module shows its input and output, so you can pinpoint where the data vanished.

A frequent gripe I’ve had is when OpenAI changed its rate‑limit header format without warning. My scenario started receiving 429 responses but the header name switched from “x-ratelimit-remaining” to “ratelimit-remaining”. Because I hard‑coded the header check, the scenario kept retrying and burned through my monthly quota. The fix was to make the header lookup case‑insensitive and to fall back to a generic 429 handler.

Another common failure point is the LinkedIn scraper. If LinkedIn updates its HTML structure, the CSS selector returns empty strings. I added a fallback that tries an alternative selector and, if both fail, flags the lead for manual review.

To avoid silent failures, I enable email notifications on any module that returns an error status. That way I know within minutes if something goes wrong.

Finally, I keep a cheap backup path: if the GPT‑4 call fails three times, the scenario inserts a generic template line (“Hi {{name}}, I noticed {{company}} is growing fast—let’s chat?”) and continues. This ensures the outreach never stops completely, even if the AI service is down.

Pricing, love, and gripes

Cost matters when you’re solo. The OpenAI API charges $0.03 per 1k tokens for GPT‑4‑turbo. My average call uses about 800 tokens, so each lead costs roughly $0.024. For 500 leads a month, that’s $12.

Make.com’s core plan is $16/mo and gives you enough operations for this volume. I find that fair—it covers the scraper, parser, API call, and email sender without hitting limits.

By contrast, Zapier’s equivalent setup would sit on their $49/mo Professional plan for the same number of tasks. I’ve found that price steep for what you get, especially when the UI feels slower for complex branching.

My concrete love is the built‑in retry logic in Make.com: you can set a module to automatically re‑run on error with exponential back‑off, which saved me during a brief OpenAI outage last month.

My concrete gripe is the lack of detailed error messages in the OpenAI Python SDK when you exceed token limits—it just returns a generic 400 with a cryptic message that forces you to inspect the response body manually.

(Which, yes, is annoying.)

For more on this exact angle, deeper coverage of AI agent platforms.

If you’d rather skip the build and deploy a working version in an afternoon, we’ve packaged this workflow as a blueprint at deepusecase.com/vault.

— The Colophon

One AI tool. Tested. Reviewed.
In your inbox every Sunday.

~3 minute read. Real outcomes from operators, not marketers.

Free. One email per Sunday. Unsubscribe in one click.