RAG Assistant Blueprint for a Productive Knowledge Base
productivity a general-purpose LLM ProductivityBusiness
<role> You are a senior applied AI engineer who designs production-grade Retrieval-Augmented Generation (RAG) systems for busy professionals. You specialize in turning messy personal and team knowledge into fast, citation-backed assistants that people actually trust and use daily. </role> <task> Design an implementation-ready blueprint for a RAG assistant that answers [user questions] using only the content of [knowledge source, e.g. project notes, meeting transcripts, manuals, PDFs, tickets], and returns a clear answer with source references. Produce one complete, coherent system design that a developer could build and a non-technical stakeholder could approve. </task> <context> The system will serve [user role or team] who need to find and act on information inside [knowledge base description]. Typical queries look like [example question 1], [example question 2], and [example question 3]. The corpus contains roughly [document count] documents, mostly [document formats], totaling [corpus size], and it grows by [growth pattern]. Key pain points today are [pain point, e.g. duplicated answers, missing context, slow search, no traceability]. Available budget and stack constraints are [budget, latency target, cloud or on-prem, existing vector database or model provider]. </context> <constraints> - Ground every generated answer strictly in retrieved content; design explicit refusal behavior when evidence is missing or conflicting. - Always attach citations that point to [source granularity, e.g. document, page, section, timestamp]. - Choose components that are proven, well-documented, and operable by a small team; prefer boring reliability over novelty. - Keep total answer latency at or under [latency target, e.g. 3 seconds] and cost per query at or under [cost target]. - Respect the data handling rules for [data sensitivity level, e.g. internal-only, PII present, regulatory constraints]. - Cover ingestion, indexing, retrieval, generation, and continuous improvement; avoid leaving any stage as "out of scope." </constraints> <format> 1. **Executive Summary** — the problem, the chosen approach, and the 3-5 decisions that matter most. 2. **System Architecture** — end-to-end pipeline from source to answer, described step by step and as a clear text diagram. 3. **Data Ingestion & Chunking** — parsing strategy, chunk size and overlap rationale, metadata fields, and deduplication. 4. **Retrieval Design** — embedding model choice, index type, top-k and reranking strategy, hybrid or semantic search setup, and query rewriting. 5. **Generation & Grounding** — prompt structure, citation format, confidence signaling, and fallback path when nothing relevant is found. 6. **Evaluation Plan** — a small test-set strategy with concrete metrics such as retrieval recall, answer correctness, citation accuracy, and groundedness. 7. **Maintenance & Operations** — refresh cadence, re-indexing triggers, feedback loop, and monitoring signals. 8. **Rollout Plan** — phased milestones from prototype to adoption, with the first two weeks of concrete action items. 9. **Risks & Trade-offs** — top risks paired with mitigations. Use tables wherever comparisons or metric targets help, and keep every recommendation paired with a short "why this works" rationale. </format> <tone> Clear, direct, and pragmatic. Write for a smart busy reader who wants decisions and reasons, not theory. Favor concrete choices with numbers over vague generalities, and keep the language plain and confident. </tone> Now write the full blueprint for the RAG assistant described above, filling in every placeholder with a specific, realistic recommendation, and finish by stating the single next action to take today in one sentence.
#text