Let the Judge Choose, Not Score
Build continuous ratings from pairwise LLM judgments, without comparing every candidate pair on every task.

AI Engineering Blog
Making AI do the thing.
Practical notes on building production RAG systems, LLM evaluations, and retrieval optimization.
Build continuous ratings from pairwise LLM judgments, without comparing every candidate pair on every task.
How human follow-up behavior reveals response quality, and why specific instruction adherence outperforms vague relevance scoring.
Using citation behavior in production RAG systems to generate labeled training data for domain-specific reranking models.
About me
Senior Staff AI Engineer at Miro, previously Reforge + Monterey.ai via acquisitions.
My boss told me the path to immortality is to write a niche blog that I'm too socially awkward to advertise. So here we are, a collection of notes on lessons learned from yelling at LLMs, professionally.
6 items
Build continuous ratings from pairwise LLM judgments, without comparing every candidate pair on every task.
How human follow-up behavior reveals response quality, and why specific instruction adherence outperforms vague relevance scoring.
A comprehensive guide to using Reciprocal Rank Fusion (RRF) to combine BM25 and semantic search results for production RAG systems.
Using citation behavior in production RAG systems to generate labeled training data for domain-specific reranking models.
Using LLM-based classification as a second pass to filter retrieval candidates when similarity thresholds fail to generalize.
Using Hypothetical Document Embeddings to bridge vocabulary gaps between user queries and specialized document corpora.