RAG Chat Template
productivity a general-purpose LLM Customer SupportAnalysis
<role>You are an expert AI engineer specializing in Retrieval-Augmented Generation (RAG) systems, vector databases, and LLM application architecture.</role> <task>Design a production-ready RAG chat template that retrieves relevant context from a knowledge base and generates accurate, well-cited responses to user queries.</task> <context>You are building a reusable RAG chat template for [target_use_case: e.g., customer support, internal documentation, research assistant, legal document analysis]. The system will ingest [document_types: e.g., PDFs, markdown files, Confluence pages, Notion databases] from [data_sources: e.g., local directory, cloud storage, API endpoints], chunk them using [chunking_strategy: e.g., recursive character splitting, semantic chunking, fixed-size with overlap], embed them with [embedding_model: e.g., text-embedding-3-large, bge-large-en-v1.5, instructor-xl], store in [vector_database: e.g., Pinecone, Weaviate, Chroma, Qdrant], and generate responses using [llm_model: e.g., GPT-4o, Claude 3.5 Sonnet, Llama 3.1 70B] with [retrieval_k: e.g., 5-10] top chunks.</context> <constraints> - Implement hybrid search (vector + keyword) for optimal retrieval - Include query rewriting/decomposition for complex questions - Add reranking step with [reranker_model: e.g., bge-reranker-large, cohere-rerank-v3] - Enforce citation format with document IDs and relevance scores - Handle edge cases: no relevant context, conflicting information, out-of-scope queries - Implement conversation memory with [memory_window: e.g., last 5 turns] - Add guardrails for hallucination detection and PII protection - Support streaming responses for better UX - Include evaluation harness with [eval_metrics: e.g., faithfulness, answer_relevancy, context_precision] - Make configuration fully parameterized via YAML/JSON config file </constraints> <format>Provide the complete template as a structured Python project with: 1. Project structure diagram 2. Configuration schema (config.yaml) 3. Core modules: ingestion, chunking, embedding, retrieval, generation, evaluation 4. Main pipeline orchestration class 5. CLI entry point with commands: ingest, query, evaluate, serve 6. FastAPI/Gradio server implementation 7. Dockerfile and docker-compose.yml 8. Comprehensive README with usage examples 9. Unit and integration tests 10. Example evaluation dataset format</format> <tone>Technical, precise, engineering-focused, and production-oriented. Use clear code comments and type hints throughout.</tone> <final_instruction>Generate the complete RAG chat template project structure with all core modules, configuration, and documentation. Begin with the project tree diagram and config.yaml schema.</final_instruction>
#text