Duck
Tutorials6 min read

Building a Relevance AI System from Scratch: A Practical Guide for Solo Operators

Samet Turan— Editor··6 min read

Learn how to build a relevance AI pipeline for content recommendation or lead scoring using open‑source tools, real prompts, and a step‑by‑step workflow.

Building a Relevance AI System from Scratch: A Practical Guide for Solo Operators

Many solopreneurs waste hours manually sorting leads or articles because they lack a way to rank relevance at scale. After reading this, you’ll be able to assemble a working relevance AI pipeline that scores items using vector embeddings and simple rules, without hiring a data scientist.

You’ll see concrete tool names, actual prompts, and a clear failure mode that trips up most beginners. Stick with me and you’ll have a repeatable process you can run in an afternoon.

Why relevance matters for solo operators

Relevance is the quiet engine behind any feed that feels personal. When you can surface the most pertinent item first, you keep attention and drive action.

For a freelancer juggling client outreach, a relevance score can turn a cold list into a warm queue. For a small agency running a blog, it can surface the next article a reader is most likely to enjoy.

The problem isn’t a lack of data; it’s the lack of a cheap, repeatable way to turn raw text into a ranked list.

Most tutorials assume you have a team of engineers and a budget for proprietary APIs. That assumption leaves solo operators stuck with manual sorting or brittle keyword matches.

What follows is a low‑cost, open‑source approach that works on a laptop and scales to a few thousand items without breaking a sweat.

What most guides get wrong about building relevance AI

Many guides start by telling you to “just call an embedding API” and then stop. They skip the messy part where you have to decide which embedding model fits your domain.

They also pretend that similarity search is a plug‑and‑play black box. In reality, the distance metric you choose (cosine vs dot product) can flip your rankings if you don’t normalize vectors.

Another common mistake is to treat the relevance score as a final answer. Most guides never show you how to calibrate that score against a business goal like click‑through or reply rate.

Finally, they ignore the operational overhead: where do you store the vectors, how do you update them when new content arrives, and what happens when the index grows beyond memory?

If you follow those guides you’ll end up with a demo that works on a handful of sentences and falls apart the moment you load real data.

How to debug when this breaks

When your relevance pipeline returns weird rankings, the first place to look is the embedding step. Check that the model you’re using actually sees the text in the language you expect.

Run a quick sanity check: embed two identical strings and verify the distance is near zero. If it’s not, something is off with the tokenizer or the model version.

Next, inspect the index. Most vector stores let you dump a few vectors and their IDs. Make sure the IDs line up with the source items you think they do.

If the rankings look reversed, verify that you’re sorting in the correct direction. Cosine similarity ranges from -1 to 1, but many libraries return distance (1‑cosine) which needs to be inverted.

Finally, watch for index staleness. If you add new documents but never re‑index, the search will never see them. A simple timestamp check on the source folder can catch this before it confuses you.

One concrete gripe: the documentation for Pinecone’s free tier assumes you already know how to generate embeddings, and the getting‑started guide skips over the step where you actually turn raw text into vectors. I spent an hour staring at a blank index because I missed that detail.

A concrete named example: setting up a vector store with Weaviate and scoring blog posts

Let’s walk through a real‑world scenario: you have a collection of 500 blog posts and you want to show the three most relevant posts at the end of each article.

First, spin up a free Weaviate sandbox cluster. The free tier gives you enough storage for a few thousand vectors and a GraphQL endpoint you can hit with curl.

Next, choose an embedding model. For English blog text, the sentence‑transformers/all-MiniLM-L6-v2 model works well and is small enough to run on a CPU.

Here’s a numbered list of the steps you’ll run in a terminal:

  1. Install the Python client: pip install weaviate-client sentence-transformers
  2. Download the model and encode your posts:
    1. from sentence_transformers import SentenceTransformer
    2. model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
    3. embeddings = model.encode(post_texts, show_progress_bar=True)
  3. Create a schema in Weaviate for a BlogPost class with a vector property:
    1. schema = {"class": "BlogPost", "vectorizer": "none", "properties": [{"name": "title", "dataType": ["string"]}, {"name": "body", "dataType": ["text"]}]}
    2. client.schema.create(schema)
  4. Upload each post with its vector:
    1. Loop over your posts, call client.data_object.create(data_object, "BlogPost", vector=emb)
  5. Run a hybrid query that combines keyword and vector search:
    1. query = {"Get": {"BlogPost": {"_additional": {"distance": true}, "limit": 3, "nearText": {"concepts": ["machine learning"]}}}}
    2. result = client.graphql.query(query)

That’s it. You now have a relevance score (the distance field) you can sort on and display.

I love how Weaviate’s GraphQL interface lets you fetch related items with a single query, cutting down on extra API calls. The hybrid search feature means you don’t have to choose between pure keyword and pure vector; you get both in one round trip.

Price mention with opinion: $29/mo for the basic Weaviate cloud tier feels fair for the features you get, but the $199/mo enterprise plan feels ridiculous unless you’re handling millions of vectors.

How do you handle low‑confidence predictions without hurting user trust?

Even a well‑tuned relevance model will sometimes return a result that feels off. Showing a low‑confidence item as the top recommendation can erode trust fast.

One practical tactic is to set a confidence threshold. If the distance (or similarity) is above a cutoff, fall back to a deterministic rule like “most recent” or “most popular”.

Another approach is to blend the relevance score with a popularity boost. For example, final_score = 0.7 * relevance + 0.3 * log(view_count). This pushes evergreen, well‑seen items up when the model is unsure.

You can also surface a “why this?” explanation. Weaviate lets you retrieve the raw vector and the nearest neighbors; showing the top matching terms from the neighbor’s title gives users a hint about the match.

If you’ve tried Zapier, you know what I mean: the platform hides the underlying math, so when something goes wrong you’re left guessing. With a self‑hosted relevance pipeline you have the knobs to tune and the visibility to see why a particular score was assigned.

Finally, keep a log of every query and the returned scores. Over time you’ll see patterns where the model consistently under‑performs on certain topics, signaling a need for fine‑tuning or a different embedding model.

Putting it all together: a simple workflow you can run today

Here’s a compact version of the steps you can copy into a bash script or a Makefile.

  1. Collect your source items (CSV, JSON, or plain text files) into a folder called ./data.
  2. Run the embedding script (embed.py) that reads each file, creates a vector, and writes a JSON lines file with id,text,vector.
  3. Feed that JSON lines file into Weaviate using the bulk import endpoint (weaviate import --schema blogpost.json --data embeddings.jsonl).
  4. Set up a cron job (or a GitHub Action) that re‑runs the embedding step whenever ./data changes and pushes the updates to Weaviate.
  5. Build a tiny front‑end (even a plain HTML page with fetch) that calls the GraphQL nearText query and displays the top three results with their distance scores.

That’s the whole loop. Once it’s running, you can forget about manual sorting and let the machine do the heavy lifting.

Adjacent reading: AI meeting tools coverage.

If you’d rather skip the build and deploy a working version in an afternoon, we’ve packaged this workflow as a blueprint at deepusecase.com/vault.

— The Colophon

One AI tool. Tested. Reviewed.
In your inbox every Sunday.

~3 minute read. Real outcomes from operators, not marketers.

Free. One email per Sunday. Unsubscribe in one click.