Open-source RAG workshop demonstrates production deployment without relying on paid APIs
Typical production RAG tutorials depend on commercial APIs, which prevents observation of actual costs or latency when using custom hardware. On August 29, AI workflow specialist Ben Auffarth led a session that concentrates on installing completely open-source RAG stacks. The presentation tackles frequent breakdowns that appear while moving from proof‑of‑concept demos to genuine applications, and it stresses a hybrid retrieval method that blends vector search with keyword search to curb LLM hallucination, particularly for queries containing terms that generate weak semantic embeddings.
Reranking receives special attention because the top‑k results from vector search often introduce irrelevant material. The workshop illustrates how applying rerankers refines precision by discarding chunks prior to their insertion into the context window.
The stack relies on RAGAS to deliver a strict assessment, moving past informal “vibe checks” such as testing three questions. This utility measures faithfulness and relevance, swapping qualitative guesses for concrete metrics. Further topics covered include:
- Guardrails built into the design phase to avoid later patches
- Benchmarking of open‑model deployments, supporting hardware capacity planning
- Hybrid search that extends beyond cosine similarity to incorporate keyword retrieval
Developers seeking to avoid vendor lock‑in receive a roadmap focused on engineering concerns, framing RAG as a technical problem rather than a prompt‑tuning exercise. Event information can be found at https://www.eventbrite.co.uk/e/the-genai-build-lab-build-production-ready-rag-on-a-budget-tickets-1994016271345?aff=rml.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Mistral local is a game changer. How much did your latency drop compared to the API? I've been running local models and the latency improvement is significant, especially when using a sophisticated hybrid retrieval approach that combines vector search with keyword search. This method often proves to be the only way to prevent LLM hallucination when users request specific terms lacking strong semantic embeddings.
This is a game changer. Which embedding model is working best with your local LLM? Combining vector search with keyword search is often the only way to prevent LLM hallucination when users request specific terms lacking strong semantic embeddings.
My token bills plummeted after switching to local models. Which hardware are you using? Most production RAG tutorials online simply wrap paid APIs, preventing real cost or latency benchmarking on your own hardware. A hands-on workshop scheduled for August 29 addresses deploying fully open-source stacks, led by Ben Auffarth, who specializes in AI workflow optimization. The technical focus targets pipeline components that typically fail when moving from demo to real-world application. Rather than dumping documents into a vector store and hoping for results, the session covers a sophisticated hybrid retrieval approach. Combining vector search with keyword search is often the only way to prevent LLM hallucination when users request specific terms lacking strong semantic embeddings. Reranking also gets detailed attention. Anyone who has built a RAG pipeline knows top-k vector search results tend to be noisy. Adding a reranker provides a practical method for improving precision by filtering chunks before they reach the context window. The inclusion of RAGAS for evaluation stands out. Too many teams "vibe check" their RAG systems — asking three questions, receiving acceptable answers, and assuming the system works. RAGAS enables deep analysis of faithfulness and relevancy, converting qualitative guesses into quantitative metrics. The workshop also covers: - Guardrails: Implementing constraints during design rather than patching them later - Benchmarking: Real-world performance data on open-model deployments for precise hardware planning - Hybrid Search: Moving beyond simple cosine similarity to incorporate keyword-based retrieval