# Personal RAG harness: Jev judges, a local model writes The code behind https://rohitghumare.com/blog/jev-personal-rag-harness/. Python 3.12, the standard library, and numpy. No framework. | File | Job | | - | - | | `ingest.py` | Reads `blog/*/index.html` and `guides/*/index.html` under a site folder, chunks by h2 section, writes table rows as `Header: value` pairs, and embeds with `nomic-embed-text`. | | `jev.py` | Minimal client for `POST https://api.typesafe.ai/v1/systemone`. Reads the key from `~/.config/typesafe/key` only. | | `route.py` | One Jev request: a Choice for intent and a Noul for "needs retrieval". | | `retrieve.py` | Cosine top 20 over the embeddings. | | `rerank.py` | One Noul per candidate passage, all in one request. Keeps the top 5. | | `generate.py` | Answers with a local qwen3 8B through ollama, citing `[c1]` to `[c5]`. | | `verify.py` | One Noul per cited sentence. Support below 0.5 escalates the answer. | | `decompose.py` | Optional. Splits a multi-part question into sub-queries. It did not help in my tests. | | `harness.py` | The CLI. Prints each stage with time, Jev tokens, and dollars. | | `evaluate.py` | Recall, MRR, McNemar, router accuracy, verify pass rate, and cost over a question file. | | `eval/questions.jsonl` | 30 questions written while reading the essays (easy). | | `eval/questions-hard.jsonl` | 30 held-out paraphrased questions (hard). | ## Setup 1. Install ollama and pull the models: `ollama pull nomic-embed-text` and `ollama pull qwen3:8b`. 2. Create the 8k-context model from the `Modelfile`: `ollama create qwen3-8b-8k -f Modelfile`. 3. Create a venv with numpy: `uv venv venv --python 3.12` then `VIRTUAL_ENV=$PWD/venv uv pip install numpy`. 4. Save your TypeSafe key in a file only you can read: `mkdir -p ~/.config/typesafe`, write the key to `~/.config/typesafe/key`, then `chmod 600 ~/.config/typesafe/key`. ## Run ``` venv/bin/python ingest.py /path/to/your/site venv/bin/python harness.py "How much memory do the weights of OpenAI's smaller open-weight model take?" venv/bin/python harness.py --no-jev "Same question, embeddings only" venv/bin/python evaluate.py --questions eval/questions-hard.jsonl --tag hard ``` `ingest.py` expects `
` pages laid out like rohitghumare.com. Point the parser at your own notes by changing `parse()` and the glob in `main()`. Without the key file, `harness.py` runs embeddings plus the local model only.