Build a Personal RAG Harness: Jev Judges, a Local Model Writes

Most RAG demos ask one model to search, choose, write, and grade its own work. This build splits those jobs. Jev makes the small typed judgments, a local model writes the words, and every stage reports its time and cost. It runs over 70 of my own essays and guides, in ten short Python files with no framework.


What we are building

You ask a question about your own writing. The harness decides whether to search, finds 20 candidate passages, and asks Jev which ones help. A local qwen3 writes an answer from the best 5 and cites each sentence. Then Jev checks every cited sentence against its passage and escalates the answer when a sentence does not hold.

The figure below is one real run from my held-out question set. The question carries a false premise on purpose. Step through it and watch the passage that answers it climb from rank 15 to rank 1.

Fig 1 · Pipeline stage tracer

One question, five stages, logged on October 2, 2026. Tap a stage. Each one shows its time, its Jev tokens, and its cost. The gold page is my essay on tokens per second, and its chunk carries a green tag.

Measured: one run of the harness on the held-out question h30 against jev-1.13.0, nomic-embed-text, and qwen3 8B on an Apple Silicon Mac. Generate time depends on your machine. The Jev stages took about half a second each in every run.

The tech stack

I have no affiliation with TypeSafe, and I paid for my own API usage. Every logged run and experiment behind this post cost about 22 cents of Jev time in total.

Each step below explains the mechanism first, then shows the code from the published files. The measured numbers come from two question sets. The first set has 29 questions and the held-out set has 30. I trimmed one question from the first set after the runs, and every number here is recomputed without it.

Step 1: Chunk by meaning, and keep tables as pairs

Retrieval returns chunks, so a chunk is the unit of evidence. A good chunk holds one idea and the words that name it. A fixed 500-token window ignores that. It cuts a definition from its example, or a number from the sentence that says what it counts.

My essays already mark their ideas with h2 sections. So the parser starts a new chunk at every h2 and keeps the section title with it. Sections over 360 words become 300-word windows with 50 words of overlap. Short sections stay whole.

Tables need one more rule. HTML flattens a table into a stream of cells, and the column names disappear. My first build turned a results table into one line of numbers. The writer then claimed "the hosted Jev API scored 21 out of 30." My essay says Laya scored 21 and Jev scored 30.

The fix writes each row as Header: value pairs. After it, the same baseline question returned "30 out of 30." The table fix also changed what the verifier saw, as Step 8 shows.

def emit_row(self, t):
    cells = t["row"]
    if t["head"] or (not t["headers"] and all(tag == "th" for tag in t["tags"])):
        t["headers"] = cells
        return
    if len(t["headers"]) == len(cells):
        line = "; ".join(f"{h}: {v}" if h else v for h, v in zip(t["headers"], cells) if v)
    else:
        line = " | ".join(cells)
    self.buf.append(f" {line}. ")

def windows(words):
    if len(words) <= MAX_SECTION:
        return [words]
    step = WINDOW - OVERLAP
    return [words[i:i + WINDOW] for i in range(0, len(words) - OVERLAP, step)]

Step 2: Embed with a bi-encoder, and know what it misses

An embedding model is a bi-encoder. It turns the question into one vector and each chunk into another, separately. Search is then a dot product between them. The model never sees the question and the chunk together.

That design is what makes search fast. You embed 836 chunks once, and each question costs one embedding plus a matrix product. It is also the limit. Each vector compresses its text before the question exists, so cosine measures shared topic. It cannot check whether a passage answers.

def embed(texts, prefix):
    body = json.dumps({"model": EMBED_MODEL, "input": [prefix + t for t in texts]}).encode()
    req = urllib.request.Request(f"{OLLAMA}/api/embed", body, {"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=600) as resp:
        vectors = np.array(json.load(resp)["embeddings"], dtype=np.float32)
    return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)

def retrieve(query, k=20):
    q = embed([query], "search_query: ")[0]
    scores = VECTORS @ q
    top = np.argsort(-scores)[:k]
    return [dict(CHUNKS[i], similarity=float(scores[i])) for i in top]

nomic-embed-text expects a prefix on each input. Documents get search_document: and questions get search_query:. Normalizing the vectors once turns cosine into a plain dot product.

The measured misses show the limit. On my first set, written while reading the essays, cosine put the gold page in the top 5 for all 25 questions that have one. On the held-out set, written in paraphrase, that fell to 0.808. Five gold pages fell out of the top 5:

Look at the cosine scores in Fig 1, too. The 20 candidates for h30 span 0.648 to 0.690. The ordering inside that band carries little signal. Bi-encoders are also known to be weak with negation. I did not test negation directly, and the false premise in h30 is the closest case in my set.

Step 3: Give Jev one small client

Jev is a System One model. It takes a state and a map of named questions, and returns a typed answer under each name. A Noul returns the probability of yes. A Choice returns a distribution over your options and a confidence. A Score returns a position on levels you describe. There is no generated text to parse.

Pricing follows the same shape. TypeSafe's models page charged $0.042 per million input tokens on October 2, 2026, and output tokens were free. So the cost of a call depends on how much state and question text you send.

URL = "https://api.typesafe.ai/v1/systemone"
KEY_FILE = Path.home() / ".config/typesafe/key"
PRICE_PER_INPUT_TOKEN = 0.042 / 1_000_000

def ask(state, questions, stage=""):
    body = json.dumps({"state": state, "model": "jev-latest", "questions": questions}).encode()
    headers = {"Authorization": f"Bearer {KEY_FILE.read_text().strip()}",
               "Content-Type": "application/json"}
    started = time.perf_counter()
    with urllib.request.urlopen(urllib.request.Request(URL, body, headers), timeout=60) as resp:
        data = json.load(resp)
    tokens = data["usage"]["input_tokens"]
    CALLS.append({"stage": stage, "model": data["model"], "input_tokens": tokens,
                  "seconds": time.perf_counter() - started,
                  "dollars": tokens * PRICE_PER_INPUT_TOKEN})
    return data["answers"]

Keep the key in a file only you can read, and load it at call time. The published file also retries HTTP 429 and honors retry-after. Log the tokens and seconds of every call. The cost model later in this post comes from that log.

Step 4: Route with a cheap typed decision

Generation is the expensive stage. A routing decision is cheap. So the harness decides what to do before it searches or writes. "Thanks" needs no retrieval. "What is the weather" needs data the notes do not hold.

One request asks two independent questions over one shared state: a Choice for intent and a Noul for "needs retrieval". They run in parallel inside the request and cannot see each other. That request cost about 620 tokens, or $0.000026, and took 0.45 to 0.60 seconds at p50.

"intent": {
    "type": "choice",
    "instructions": "What should the assistant do with `message`, given that it can search `notes`?",
    "criteria": {
        "answer_from_my_notes": "A question about AI models, APIs, agents, tools, benchmarks, prices, protocols, or hardware that the notes could cover, even if the topic is recent.",
        "small_talk": "A greeting, thanks, or chit-chat that needs no information.",
        "needs_live_web": "Needs real-time data that changes by the hour or day and has nothing to do with the notes: weather, stock or crypto prices, sports scores, traffic.",
        "ambiguous_ask_back": "Too vague to answer without asking the user what they mean.",
    },
},

retrieve_needed = r["intent"] == "answer_from_my_notes" or r["confidence"] < 0.5

The last line carries the policy. My first rule trusted the intent label alone and sent 5 answerable questions away. It made the right retrieve decision on 23 of 29. The current rule also retrieves when the intent confidence is below 0.5. That matches the low-confidence band in TypeSafe's confidence docs.

The costs are uneven, so the gate leans toward search. A needless search costs about half a second. A wrong refusal loses the answer. The new rule scored 27 of 29 on the first set, where I tuned it. On the held-out set it scored 29 of 30, and it handled all 4 out-of-scope messages correctly.

The "needs retrieval" Noul did not help. It read 0.06 on "Where does Claude Code put CLAUDE.md", which my notes answer directly. The harness logs it and ignores it.

Step 5: Rerank with joint judgments

A reranker fixes what the bi-encoder cannot see. It reads the question and one passage together and judges that pair. In retrieval terms, this is the cross-encoder role. TypeSafe has not published Jev's architecture. The point is the input: each judgment sees both texts at once.

Each candidate gets one Noul: does this passage help answer the query? The false criterion names the exact failure of cosine search, a passage that "only shares words with the query." All 20 questions go into one request over one shared state.

def question(key):
    return {
        "type": "noul",
        "instructions": f"Does the passage `passages.{key}` help answer `query`?",
        "criteria": {"true": "The passage states facts, numbers, or claims that directly answer the query.",
                     "false": "The passage is off topic or only shares words with the query."},
    }

def rerank(query, candidates, keep=5):
    scored = []
    for batch in batches(query, candidates):
        state = {"query": query, "passages": {k: passage(c) for k, c in batch.items()}}
        answers = ask(state, {k: question(k) for k in batch}, stage="rerank")
        scored += [dict(c, relevance=answers[k]["noul"]) for k, c in batch.items()]
    scored.sort(key=lambda c: (-c["relevance"], -c["similarity"]))
    return scored[:keep]

Batching is the main Jev-specific choice here. The query and the instructions travel once, and the 20 judgments run in parallel. TypeSafe's parallel questions cookbook measured 13 questions in one call as 12.2 times cheaper and 10 times faster than 13 calls. In my runs, one rerank request took 0.55 to 0.77 seconds at p50 and averaged 8,697 tokens. A helper, batches(), splits the list only when the payload would pass the 32k state budget.

On the held-out set, rerank lifted gold-page recall@5 from 0.808 to 0.962 and recall@1 from 0.654 to 0.808. On the first set, the right page was already in the top 5, and rerank moved the right section to the top. Section recall@1 rose from 0.636 to 0.864. Across 51 questions with a gold page, rerank never pushed a gold page out of the top 5.

Step 6: Gate on probabilities, and keep the scores

A rank says which passage scored higher. It says nothing about whether either one is good. A calibrated probability does. If the model is calibrated, 0.95 should mean the passage helps about 95 times in 100. TypeSafe trains Jev for calibrated decisions, and my Laya and Jev test measured an expected calibration error of 0.023 on 30 tickets. Calibration lets one threshold mean the same thing across every query.

The two traces I logged show why that matters. For h30, seven passages scored 0.50 or more. For the gpt-oss question h17, only two did. A fixed top 5 sends three weak passages in the first case and drops two useful ones in the second.

The harness gates on probabilities at three points. The router searches when intent confidence is below 0.5. Rerank sorts by the Noul. Verify escalates any sentence below 0.5. TypeSafe's confidence docs suggest setting each threshold by the cost of being wrong, so a destructive action deserves a higher bar than a read.

Keep every score you paid for. The harness stores each passage's relevance and each sentence's support. You can then change a threshold, a weight, or the number of passages kept without calling Jev again. TypeSafe's composite scoring pattern describes the same idea. Fig 3 below re-gates stored scores with no new calls.

This turns into a cascade. Jev answers the cheap typed questions. Only the uncertain or failing cases go to a stronger model or a person. TypeSafe's extraction cascade cookbook escalates when any per-field flag passes 0.7, so one strong signal is enough.

Step 7: Generate only from cited chunks

The writer is the one stage that can invent. So it gets the least freedom. It sees only the 5 kept passages, numbered [c1] to [c5], and four rules. Use only the notes. Cite every fact. Reply with one fixed sentence when the notes lack the answer. Keep it to five sentences.

SYSTEM = """You answer questions using only the numbered notes below.
Rules:
- Use only facts stated in the notes. Never add outside knowledge.
- End every sentence that states a fact with the id of the note it came from, like [c2].
- If the notes do not contain the answer, reply exactly: I don't know from my notes.
- Answer in at most five sentences."""

MODEL = "qwen3-8b-8k"
body = json.dumps({"model": MODEL, "messages": messages, "temperature": 0,
                   "reasoning_effort": "none"}).encode()

Citations make the next step possible. Every factual sentence names its source, so a checker can test the sentence against that one passage.

Two ollama details cost me time. The OpenAI-compatible endpoint ignored num_ctx, so I built qwen3-8b-8k from a two-line Modelfile. On this ollama build, only reasoning_effort: "none" turned qwen3's thinking off.

Step 8: Verify every claim against its source

Verification is natural language inference. The cited passage is the premise. The sentence is the hypothesis. The question is whether the premise entails it. Jev asks that as one Noul per sentence, and all sentences go into one request.

questions = {
    f"s{i}": {
        "type": "noul",
        "instructions": f"Is every fact in `sentences.s{i}` stated in "
                        + " or ".join(f"`notes.{cid}`" for cid in c["cites"]) + "?",
        "criteria": {"true": "The cited note states the facts, names, and numbers in the sentence.",
                     "false": "The sentence adds or changes a fact, name, or number the cited note does not state."},
    }
    for i, c in enumerate(claims)
}
answers = ask(state, questions, stage="verify")
for i, claim in enumerate(claims):
    claim["support"] = answers[f"s{i}"]["noul"]
weak = [c for c in claims if c["support"] < THRESHOLD]

A checker can only judge the text it reads. In the first build, verify scored the wrong "21 out of 30" sentence 0.88 and passed it. It read the same flattened table the writer read. After the table fix, it scored that wrong sentence 0.03 and the right one 0.92.

On the first set, verify passed 21 of 24 answers. On the held-out set it passed only 14 of 25. I read every escalated answer by hand. Of the 11 held-out flags, 6 were real errors, 4 were false alarms on correct answers, and 1 was unclear.

Fig 3 · Verify threshold board

Each dot is one verified answer, placed at the support score of its weakest sentence. Color comes from my hand review of the escalated answers. Move the threshold to see what a stricter or looser gate catches. Tap a dot to read it.

real errorfalse alarmunclearpassed, not hand-judged
0.50
Tap a dot to read the weakest sentence.

Measured on October 2, 2026, jev-1.13.0. I judged only the answers verify escalated at 0.5. Answers that passed were not judged by hand, so this board cannot show what verify missed.

The real catches were useful. Answer h29 invented an earlier price of "$1.00". Answer h22 named the wrong guide. On the first set, q06 cited an introduction that lacked the number it stated.

The false alarms clustered on numbers and close paraphrase. Answer h16 gave a correct table value, 147.92 GB, and verify still doubted it. TypeSafe's jev-1.13 jaggedness page says the model "struggles with tasks that require numeric precision." The trace in Fig 1 shows another false alarm. Its last sentence scored 0.25, though the passage says a bandwidth-heavy card matters more than a compute-heavy one.

So pair verify with a check that needs no model. Every number in the answer must appear in the text of a chunk it cites. On the held-out set, answers from the Jev pipeline passed that check 7 of 7 times. Ship an answer when both checks pass, and review it when either fails.

Step 9: Wire the loop and print every stage

The harness is a plain function. It times each stage, attributes Jev calls to it, and adds up tokens and dollars. A --no-jev flag runs embeddings plus the writer only, which is the baseline for every comparison in this post.

$ venv/bin/python harness.py "Since running models at home is limited by raw compute, should I just buy the card with the most TFLOPS?"
route         489 ms     629 tok  $0.000026
retrieve       68 ms
rerank        584 ms    7187 tok  $0.000302
generate     8584 ms
verify        383 ms    2416 tok  $0.000101

route: answer_from_my_notes (confidence 1.00), needs_retrieval 0.09
  [c1] jev 0.95  Tokens per Second Is a Memory Bandwidth Number / What this predicts
  [c2] jev 0.89  How to Run Kimi K3 Locally / Can you run it?
  [c3] jev 0.89  How to Run DeepSeek R1 and V3.2 Locally / The full 671B with offload
  [c4] jev 0.86  How Much RAM, and Which Mac or GPU / The mixture-of-experts exception
  [c5] jev 0.56  How to Run GLM-5.3-Flash Locally / What fits where

verify: ESCALATE, 5 cited sentences checked
  0.51  No, raw compute is not the limiting factor for running models at home; memory bandwidth and capacity
  0.91  A faster GPU may not compensate if your model exceeds VRAM and spills to system RAM, which is read a
  0.95  For mixture-of-experts models, the speed is determined by the active experts, not the total paramete
  0.82  Matching your combined VRAM and system RAM to the quant file size ensures better throughput.
  0.25  Buying the card with the most TFLOPS might not be the best choice if your system lacks sufficient me

total 10108 ms, Jev 10232 tokens, $0.000430

The gold answer in my question file says memory bandwidth sets speed at batch size one. The pipeline rejected the premise and cited the right essay, which cosine had ranked 15th. The escalation is the false alarm from Step 8, and a reviewer would clear it in seconds.

Measure it on questions you did not copy

One good run proves little. I wrote two question sets, marked a gold page for each answerable question, and ran each set with and without Jev.

The first set was too easy. I wrote it while reading the essays, so the questions reused the essays' words. Cosine already found every gold page. Rerank only moved the right section of the right page to the top: 5 wins and 0 losses at section recall@1, McNemar p = 0.0625.

The held-out set was harder. I wrote 30 new questions from page titles and memory, in plain paraphrase. I froze the file before any run and recorded its hash. This is the set that showed the rerank gain.

Fig 2 · Rank-shift ladder

Each line is one question. The left end is the gold rank from cosine search. The right end is the gold rank after Jev rerank. Switch the question set and the level. Tap a question below to trace its line.

rank improvedrank unchangedrank worse
Tap a question to read it.

Measured on October 2, 2026. Same 20 candidates in both columns, so rerank only reorders. Ranks past 20 mean the gold section was not among the candidates. McNemar is the exact two-sided test on discordant pairs at the chosen cutoff.

None of these tests reach p < 0.05. On the held-out set, page recall@5 had 4 rerank-only wins and 0 losses, p = 0.125. Thirty questions cannot prove a small effect. The direction held on both sets with zero page-level losses, and I would act on that while a bigger set confirms it.

The rescued questions show the mechanism from Step 2. Question h09 asked where "paid API passwords" should live, and its gold page moved from rank 17 to 3. Question h30 moved from 15 to 1. Each question shared ideas with its gold page and few words.

Why decomposition did not help

Decomposition splits a multi-part question into sub-queries, searches each one, and reranks the union. It helps when one query vector cannot sit near two topics at once. I tried it with qwen3 writing up to three sub-queries.

The splits looked sensible. The results matched plain rerank. The first set had 0 discordant questions at every level. The held-out set had 1 or 2, all with p = 1.0. On cross-essay questions, the five kept passages held every gold page for 3 of 6 without Jev and 2 of 6 with decomposition.

The reason is in the numbers. Rerank already sees 20 candidates, and every held-out gold page that cosine ranked low still sat inside that window. The one cross-essay miss I traced was a recall failure: a guide that never entered the top 20. Decomposition also cost 4.6 to 16 seconds of local model time per question at p50. It stays in the code behind a flag, off by default.

A cost and latency model per stage

StageRuns onJev tokens per callCost per callp50 time
routeJev, 2 questions623$0.0000260.45 to 0.60 s
retrieveollama embedding and numpy0local0.03 to 0.35 s
rerankJev, 20 questions8,697$0.0003650.55 to 0.77 s
generatelocal qwen3 8B0local1.1 to 36 s, machine bound
verifyJev, 1 question per sentence1,970$0.0000830.47 to 0.61 s

Token counts are means from the plain pipeline run on the first set, where retrieval happened on 23 of 29 questions. Averaged over all 29, a question used 9,083 Jev tokens. At $0.042 per million, that is $0.000381 per question, or about $0.38 per thousand questions.

Rerank is 76% of that spend. The reason is arithmetic. 8,697 tokens over 20 passages is about 435 tokens per passage, including the shared query and question text. Route and verify read short texts, so they stay cheap.

Latency splits the same way. The three Jev stages added 1.5 to 2.0 seconds per question at p50. The local writer dominated wall time. Per-question generate time ranged from 0.4 to 91 seconds as other work shared the GPU.

The models page listed limits of 100,000 tokens per second and 40 requests per second on October 2, 2026. At about 9,100 tokens per question, the token limit allows about 11 questions per second. Each question makes 3 requests, so the request limit allows about 13. Tokens bind first.

Optimise it

Rerank holds three quarters of the spend, so it is where tuning pays. I ran 13 configurations of the first stage and the rerank step on October 2, 2026. Every configuration used the same 51 questions with a gold page and 41 with a gold section, from both sets. Each one ran once, and those experiments cost $0.156 of Jev time.

ConfigurationJev tokens per questionCost per questionPage recall@1 / @5Section recall@1 / @5
Embeddings only, no Jev0$00.804 / 0.9020.512 / 0.805
Hybrid BM25 plus embeddings, no Jev0$00.902 / 0.9610.537 / 0.780
K = 5, full passage, Noul2,322$0.0000980.863 / 0.9020.707 / 0.805
K = 10, full passage, Noul4,447$0.0001870.882 / 0.9410.659 / 0.805
K = 20, full passage, Noul (default)8,793$0.0003690.922 / 0.9800.732 / 0.854
K = 30, full passage, Noul12,948$0.0005440.902 / 1.0000.732 / 0.927
K = 20, first 120 words, Noul5,309$0.0002230.902 / 0.9800.683 / 0.829
K = 20, first 60 words, Noul3,648$0.0001530.863 / 0.9800.634 / 0.756
K = 20, full passage, Score9,573$0.0004020.902 / 0.9800.683 / 0.854
BM25 K = 20, Noul9,036$0.0003800.941 / 0.9610.634 / 0.756
Hybrid K = 20, Noul8,979$0.0003770.941 / 0.9800.659 / 0.829

Keep K = 20, full passages, and a Noul as the default. No configuration beat it on quality at a lower cost. The numbers in this table come from a fresh run, so they differ slightly from the logged runs above.

Use K = 5 as a budget mode. It costs $0.000098 per question, 73% less, and its p50 request time was 461 ms. Page recall@5 falls from 0.980 to 0.902. The loss is 4 held-out questions whose gold page ranked 6 to 20 on embeddings. A reranker cannot recover a passage it never sees. On page recall@5, the paired McNemar test gave 4 wins for K = 20 and 0 for K = 5. At p = 0.125, the gap is not significant.

Candidate count is a cost line. Each candidate adds about 430 tokens. K = 30 raised section recall@5 from 0.854 to 0.927 for 47% more cost. K = 10 was worse than K = 5 at section recall@1, because extra near-miss passages pushed gold sections down.

Shorter passages lost quality. Sending the first 120 words saved 40% of tokens and dropped section recall@1 from 0.732 to 0.683. The first 60 words dropped it to 0.634. Jev needs the sentence that answers, and it is often not in the first lines.

A Score did not beat a Noul. Four described levels, ranked by expected level, cost 9% more tokens. Section recall@1 fell to 0.683 and page recall@1 to 0.902. I tried only one wording for the levels.

Keyword search helps pages and hurts sections. BM25 and hybrid reciprocal rank fusion both raised page recall@1 after rerank to 0.941. They lowered section recall@1 to 0.634 and 0.659. On the held-out set, BM25 plus rerank reached only 0.474 section recall@5, because paraphrased questions share few exact words with their answers.

The cascade works, but it fires rarely. The rule escalates a question when the top reranked passage scores below a threshold. At 0.9, it escalated 13.7% of page-level questions, and recall@1 on the kept questions rose to 0.955. At section level it escalated 4.9%, and both escalated questions were section misses. At 0.5 it escalated nothing, because Jev is confident on most top passages, including some wrong ones.

Batch every judgment into one request. On 5 held-out questions, one request with 20 parallel Nouls used 8,155 tokens per question. Twenty single requests used 13,805. One request took 0.54 seconds per question, and 20 sequential requests took 9.4 seconds.

The answers were not identical across the two modes. Top 1 matched on 3 of 5 questions. Probabilities for the same passage moved by up to 0.42, while repeated batched runs moved by 0.01 to 0.07. In a batch, each judgment sees the other passages in the shared state, so the batch contents shape the scores. Keep the batch fixed when you compare two runs. Five questions cannot show which mode ranks better.

Every configuration ran once. Repeated runs move Jev probabilities by up to 0.07, which can flip a close top 1, so differences of one or two questions are noise. I also chose among 13 configurations on the same 51 questions, which overfits a little. A third question set would confirm the choice.

One measured warning applies to a cosine floor, a cutoff that drops low-similarity candidates before rerank. I did not test a floor. In the h30 trace, the passage that answered scored 0.652, only 0.004 above the lowest candidate. A floor that trims the bottom of the list would have cut it.

Failure modes and fixes

FailureWhat I sawFix
Flattened tables"Jev scored 21 of 30", passed by verify at 0.88Write rows as Header: value pairs at ingest
Paraphrase hides the gold pageHeld-out recall@5 of 0.808 with cosine aloneRerank 20 candidates with one Noul each
Router refuses answerable questions5 sent away under the label-only ruleRetrieve whenever intent confidence is below 0.5
Gold page never retrievedA cross-essay question missed one guide in the top 20Rerank cannot fix recall. Widen the first stage, then measure
Verify doubts correct numbers4 false alarms on the held-out set, one a table valuePair verify with a number-in-chunk check
Easy questions hide the gainPage recall@5 was 1.000 with or without rerankWrite a frozen, paraphrased held-out set
Slow local generationA one-line reply took 14.5 s under GPU contentionSet a small context, turn thinking off, and give the model the GPU

Run it on your notes

Start with your own writing. Change parse() for your file format. Write 30 questions you cannot answer from memory, and freeze them before the first run. Then run the evaluation with and without Jev, and keep only the stages that move a number.

The full code is published next to this post. Start with the README.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. I have no relationship with TypeSafe AI. Every number here comes from my runs on October 2, 2026: ollama 0.31.1, nomic-embed-text, qwen3 8B, and jev-1.13.0. The reranking cookbook and limits come from docs.typesafe.ai on the same date.

Related: Laya and Jev · Evals and Benchmarks · More posts · X