{"id": "q01", "kind": "number", "intent": "answer_from_my_notes", "question": "What share of Cursor's user traffic did OpenAI models serve, according to Cursor's CEO?", "gold": ["https://rohitghumare.com/blog/openai-cuts-cursor/"], "section": "Introduction", "answer": "About 5%.", "why_hard": "Cursor appears in several harness essays; the number sits in a quote."} {"id": "q02", "kind": "number", "intent": "answer_from_my_notes", "question": "What did Artificial Analysis measure as the cost per task for Gemini 3.8 Flash, compared with 3.7 Flash?", "gold": ["https://rohitghumare.com/blog/flash-cadence/"], "section": "Introduction", "answer": "$0.58 per task, up from $0.40.", "why_hard": "Two adjacent version numbers and two dollar figures; embeddings blur numbers."} {"id": "q03", "kind": "factual", "intent": "answer_from_my_notes", "question": "Which two harness settings moved OpenAI's ARC-AGI-3 score from 13.3 to 38.3 percent?", "gold": ["https://rohitghumare.com/blog/inside-the-codex-harness/"], "section": "Introduction", "answer": "Retain reasoning and enable compaction.", "why_hard": "13.3 also appears in the Muse Glimmer guide; compaction is discussed in several harness essays."} {"id": "q05", "kind": "number", "intent": "answer_from_my_notes", "question": "How large is the DeepSeek-V4.1-Flash download, and how much of it is the lookup table?", "gold": ["https://rohitghumare.com/guides/deepseek-v4-1-flash/"], "section": "Where the 510 GB goes", "answer": "510 GB, of which 203 GB is the Engram lookup table.", "why_hard": "The older DeepSeek V4 guide also mentions V4.1-Flash and 203 GB in a forward pointer."} {"id": "q06", "kind": "number", "intent": "answer_from_my_notes", "question": "How many messages did the human operator send during the RSA-260 factorization run?", "gold": ["https://rohitghumare.com/blog/rsa-260-agent-fleet/"], "section": "The steering log", "answer": "3,328 messages.", "why_hard": "Easy page match, but the number lives in a deep section."} {"id": "q07", "kind": "factual", "intent": "answer_from_my_notes", "question": "Which two things did the July 28, 2026 MCP release remove from the wire?", "gold": ["https://rohitghumare.com/blog/stateless-mcp/"], "section": "Introduction", "answer": "The initialize handshake and the Mcp-Session-Id header.", "why_hard": "The MCP roadmap essay recaps the same release and competes for the top slot."} {"id": "q08", "kind": "factual", "intent": "answer_from_my_notes", "question": "In which year and city does the graph engineering essay place the birth of the discipline's shape?", "gold": ["https://rohitghumare.com/blog/graph-engineering/"], "section": "Introduction", "answer": "1736, Konigsberg (Euler).", "why_hard": "Historical detail inside an AI essay; low lexical overlap with 'birth of the discipline'."} {"id": "q09", "kind": "factual", "intent": "answer_from_my_notes", "question": "At batch size one, what determines tokens per second for local generation?", "gold": ["https://rohitghumare.com/blog/tokens-per-second-memory-bandwidth/"], "section": "Introduction", "answer": "Memory bandwidth divided by the bytes read per token, mostly the weights.", "why_hard": "The hardware guide, inference engines essay, and model guides all discuss speed."} {"id": "q10", "kind": "number", "intent": "answer_from_my_notes", "question": "By default, how many tokens does each vLLM KV cache block hold?", "gold": ["https://rohitghumare.com/guides/inside-vllm/"], "section": "A memory problem wearing a throughput costume", "answer": "16 tokens.", "why_hard": "Deep section; many model guides mention vLLM and KV cache."} {"id": "q11", "kind": "number", "intent": "answer_from_my_notes", "question": "Which DeepSeek-V4-Flash quant is bit-for-bit lossless, and how big is it?", "gold": ["https://rohitghumare.com/guides/deepseek-v4/"], "section": "Pick your quant", "answer": "The Q8 GGUF, at 162 GB.", "why_hard": "162 GB also appears in the older DeepSeek R1/V3.2 guide; three DeepSeek guides compete."} {"id": "q12", "kind": "factual", "intent": "answer_from_my_notes", "question": "What name did GLM-5.3-Flash go by on OpenRouter before z.ai revealed it?", "gold": ["https://rohitghumare.com/guides/glm-5-3-flash/"], "section": "Introduction", "answer": "ox-alpha.", "why_hard": "The older GLM guide competes; the answer is a rare token."} {"id": "q13", "kind": "factual", "intent": "answer_from_my_notes", "question": "Who does TechCrunch name as the founder behind Jev?", "gold": ["https://rohitghumare.com/blog/laya-jev-system-1-models/"], "section": "Jev: one set of weights behind an API", "answer": "Diogo Almeida, a former OpenAI researcher.", "why_hard": "Single sentence in a long essay with 24 chunks from the same page."} {"id": "q14", "kind": "factual", "intent": "answer_from_my_notes", "question": "Where does Claude Code put CLAUDE.md in each API request?", "gold": ["https://rohitghumare.com/blog/inside-the-claude-code-harness/"], "section": "Three layers, ordered by volatility", "answer": "As a user message after the system prompt, not inside the system prompt.", "why_hard": "CLAUDE.md is discussed in the AGENTS.md essays and the Claude levels essay too."} {"id": "q15", "kind": "factual", "intent": "answer_from_my_notes", "question": "At what path does an A2A agent publish its agent card?", "gold": ["https://rohitghumare.com/blog/a2a-protocol/"], "section": "Discovery is one file at a fixed path", "answer": "/.well-known/agent-card.json (renamed from agent.json in 0.3).", "why_hard": "The Muse essay also publishes keys at a well-known path."} {"id": "q16", "kind": "number", "intent": "answer_from_my_notes", "question": "How many AGENTS.md files does the main OpenAI repo carry?", "gold": ["https://rohitghumare.com/blog/agents-md-best-practices/"], "section": "The number nobody budgets: it costs tokens on every request", "answer": "88.", "why_hard": "Two AGENTS.md essays compete; 88 also appears as vLLM's 88,000 stars."} {"id": "q17", "kind": "factual", "intent": "answer_from_my_notes", "question": "What is the standard fix for the dual-write trap?", "gold": ["https://rohitghumare.com/blog/dual-write-trap/"], "section": "The fix: the outbox pattern", "answer": "The outbox pattern: write the event to an outbox table in the same transaction, then relay it.", "why_hard": "Easy page; tests whether the fix section beats the intro."} {"id": "q18", "kind": "factual", "intent": "answer_from_my_notes", "question": "When did Windsurf become Devin Desktop?", "gold": ["https://rohitghumare.com/blog/inside-the-devin-harness/"], "section": "Four surfaces, two loops", "answer": "June 2, 2026.", "why_hard": "Date lookup in a deep section; dates are weak for embeddings."} {"id": "q19", "kind": "factual", "intent": "answer_from_my_notes", "question": "How much unified memory can a Mac Studio have, and what can it run according to the hardware guide?", "gold": ["https://rohitghumare.com/guides/hardware/"], "section": "Apple Silicon: Mini, Pro, Studio", "answer": "Up to about 512 GB; large MoE and 600B-plus models at dynamic quants.", "why_hard": "Every model guide has a Mac section that competes."} {"id": "q20", "kind": "number", "intent": "answer_from_my_notes", "question": "How many levels does the Claude levels essay count between a chat box and a fleet of agents?", "gold": ["https://rohitghumare.com/blog/claude-code-levels/"], "section": "Introduction", "answer": "Nine.", "why_hard": "Spelled-out number; other Claude essays compete."} {"id": "q21", "kind": "cross", "intent": "answer_from_my_notes", "question": "Which two model guides describe a model whose biggest tensor is a lookup table rather than a weight matrix?", "gold": ["https://rohitghumare.com/guides/qwen3-8-flash-next/", "https://rohitghumare.com/guides/deepseek-v4-1-flash/"], "section": "", "answer": "Qwen3.8-Flash-Next and DeepSeek-V4.1-Flash.", "why_hard": "Needs two pages in the top 5; one page uses 'n-gram embeddings', the other 'Engram'."} {"id": "q22", "kind": "cross", "intent": "answer_from_my_notes", "question": "What do the Codex harness and Pi harness essays each treat as the core machinery around the prompt?", "gold": ["https://rohitghumare.com/blog/inside-the-codex-harness/", "https://rohitghumare.com/blog/inside-the-pi-harness/"], "section": "", "answer": "Codex: compaction (and retained reasoning) as an API primitive. Pi: the prompt cache, shown to the user as a meter.", "why_hard": "Two pages, and the 'harness' essays (coding-agent-harnesses, harness-engineering) are strong distractors."} {"id": "q23", "kind": "cross", "intent": "answer_from_my_notes", "question": "Why did Max 20x subscribers burn a five-hour session in about twenty minutes on Fable 5.1, and how does that connect to Claude having two meters?", "gold": ["https://rohitghumare.com/blog/fable-burned-the-meter/", "https://rohitghumare.com/blog/claude-two-meters/"], "section": "", "answer": "The 75% cache-read cut was to the API list price; the plan meter was never told about it. Claude has two meters (five-hour session and weekly), and the 20x multiplier is printed on only one.", "why_hard": "Synthesis across two essays plus the Fable 5.1 launch essay as a distractor."} {"id": "q24", "kind": "oos", "intent": "small_talk", "question": "Hey! Thanks, that was really helpful.", "gold": [], "section": "", "answer": "", "why_hard": "Router must skip retrieval."} {"id": "q25", "kind": "oos", "intent": "small_talk", "question": "Good morning, how are you doing today?", "gold": [], "section": "", "answer": "", "why_hard": "Router must skip retrieval."} {"id": "q26", "kind": "oos", "intent": "needs_live_web", "question": "What's the weather in London right now?", "gold": [], "section": "", "answer": "", "why_hard": "Needs live data; notes cannot answer it."} {"id": "q27", "kind": "oos", "intent": "ambiguous_ask_back", "question": "Can you tell me more about that thing?", "gold": [], "section": "", "answer": "", "why_hard": "No referent; the right move is to ask back."} {"id": "q28", "kind": "trick", "intent": "answer_from_my_notes", "question": "Is it true that the hosted Jev API got fewer than half of the 30 test tickets right?", "gold": ["https://rohitghumare.com/blog/laya-jev-system-1-models/"], "section": "Same 30 tickets, both models, live", "answer": "No. Jev got all 30 right; the open Laya checkpoint was the one that was often confidently wrong.", "why_hard": "False premise; Laya's lower score sits next to Jev's in the same chunks."} {"id": "q29", "kind": "trick", "intent": "answer_from_my_notes", "question": "Which Gemma 3 sizes cannot read images?", "gold": ["https://rohitghumare.com/guides/gemma-3/"], "section": "Images and vision", "answer": "The 1B; the 4B and larger models can see images.", "why_hard": "Negation; the text states the positive case only."} {"id": "q30", "kind": "trick", "intent": "answer_from_my_notes", "question": "Did the July 2026 MCP release add a new session handshake to the protocol?", "gold": ["https://rohitghumare.com/blog/stateless-mcp/"], "section": "Introduction", "answer": "No. It removed the initialize handshake and the Mcp-Session-Id header.", "why_hard": "False premise phrased as the opposite of the fact."}